Task processing method based on double control units, neural network processor and equipment

Through the dual control unit architecture and the arbitrator to coordinate resource use, the problem of insufficient utilization of NPU resources is solved, and the computing efficiency and resource utilization are improved.

CN120471122APending Publication Date: 2025-08-12JIANGSU TSINGMICRO INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510342966.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Due to the data dependence between tasks, some resources cannot be fully utilized, affecting the computing efficiency.

Method used

Using a dual control unit architecture, initialization registers are configured to the first and second control units respectively, and resource usage is coordinated through an arbitrator to ensure that the two control units can perform different neural network tasks at the same time.

Benefits of technology

It significantly improves the working efficiency and resource utilization of NPUs, and achieves more efficient allocation and collaborative use of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471122A_ABST
    Figure CN120471122A_ABST
Patent Text Reader

Abstract

The invention relates to the field of neural networks, and provides a task processing method based on double control units, a neural network processor and equipment. The method comprises the following steps: configuring initialization registers for a first control unit and a second control unit respectively; the configuration of the initialization register for the second control unit at least comprises the following steps: configuring a feature map storage offset address in an offset address register of the second control unit; controlling the first control unit to automatically execute the first neural network task; controlling a second control unit to automatically execute a second neural network task; wherein the second neural network task is different from the first neural network task. According to the invention, two control units are arranged, so that the NPU is endowed with the capability of scheduling two neural network instructions at the same time. The dual control unit architecture aims to efficiently allocate and cooperatively use computing resources, so that the working efficiency and resource utilization rate of the NPU are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of neural network processors, and in particular to a task processing method based on dual control units, a neural network processor, and a device. Background Art

[0002] The rapid development of the internet, the Internet of Things, and social media has generated vast amounts of data, providing a rich source of training material for AI and deep learning. Deep learning models have achieved remarkable results in areas such as image recognition, natural language processing, and speech recognition, driving the widespread application of AI. AI applications inevitably rely on processors to complete the associated computing tasks.

[0003] Currently, NPUs hold an unshakable position in neural network applications thanks to their efficient computing capabilities, occupying an absolute market share. However, although the original design of NPUs was intended to achieve parallel computing and efficient resource utilization, in actual use, because the system can only process one neural network task at a time, some resources will inevitably not be fully utilized.

[0004] Although the NPU itself has parallel capabilities, when each unit executes different tasks, due to certain data dependencies between tasks, these tasks cannot be executed completely in parallel. In this case, some units will inevitably be idle and unable to fully realize their computing potential.

[0005] This data dependency is one of the main reasons for the limited NPU resource utilization, as it hinders the complete parallel processing of tasks, thus affecting the overall computing efficiency. Therefore, although the NPU has significant advantages in neural network computing, in practical applications, the problem of insufficient resource utilization still needs to be solved to further improve system performance. Summary of the Invention

[0006] In view of this, the embodiments of the present disclosure provide a task processing method, a neural network processor and a device based on a dual control unit to solve the problem in the prior art that some units of the NPU are idle and cannot fully utilize their computing potential.

[0007] According to a first aspect of an embodiment of the present disclosure, a task processing method based on a dual control unit is provided, comprising: configuring initialization registers to a first control unit and a second control unit respectively; wherein the first control unit and the second control unit are control units in the same neural network processor; configuring the initialization register to the second control unit comprises at least: configuring a feature map storage offset address in an offset address register of the second control unit; wherein the above-mentioned feature map storage offset address is used to determine the capacity of the feature map buffer allocated to the first control unit and the second control unit; controlling the first control unit to automatically execute a first neural network task; controlling the second control unit to automatically execute a second neural network task; wherein the second neural network task is different from the first neural network task.

[0008] According to a second aspect of an embodiment of the present disclosure, a neural network processor is provided, comprising: a first control unit, a second control unit, a functional module, and an arbitrator; wherein: the first control unit is used to execute a first neural network task; the second control unit is used to execute a second neural network task; the functional module is used to perform data transmission and calculation; and the arbitrator is used to coordinate configuration requests of the first control unit and the second control unit.

[0009] According to a third aspect of the embodiments of the present disclosure, a computer system is provided, characterized by a neural network processor of the computer system.

[0010] In a fourth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0011] Compared with the prior art, the beneficial effects of the embodiments of the present disclosure are as follows: first, initialization registers are configured for the first control unit and the second control unit respectively; wherein the first control unit and the second control unit are control units in the same neural network processor; second, configuring the initialization register for the second control unit at least includes: configuring the feature map storage offset address in the offset address register of the second control unit; wherein the above-mentioned feature map storage offset address is used to determine the capacity of the feature map buffer allocated to the first control unit and the second control unit; then, the first control unit is controlled to automatically execute the first neural network task; finally, the second control unit is controlled to automatically execute the second neural network task; wherein, the second neural network task is different from the first neural network task. The present disclosure gives the NPU the ability to schedule two neural network instructions at the same time by setting two control units. This dual control unit architecture is designed to efficiently allocate and coordinate the use of computing resources, thereby significantly improving the work efficiency and resource utilization of the NPU. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 is a flow chart of some embodiments of a dual control unit-based task processing method according to the present disclosure;

[0014] Figure 2 is a basic NPU architecture diagram of the dual control unit-based task processing method according to the present disclosure;

[0015] Figure 3 is a dual control unit NPU architecture diagram of a dual control unit-based task processing method according to the present disclosure;

[0016] Figure 4 1 is a schematic structural diagram of some embodiments of a task processing device based on a dual control unit according to the present disclosure. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0018] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0023] Figure 1 This is a schematic diagram of an application scenario of a task processing method based on a dual control unit according to some embodiments of the present disclosure.

[0024] Figure 1 Flowcharts of some embodiments of the task processing method based on dual control units according to the present disclosure. Figure 1 As shown, the task processing method based on the dual control unit includes:

[0025] Step S101, configuring initialization registers for a first control unit and a second control unit respectively; wherein the first control unit and the second control unit are control units in the same neural network processor.

[0026] In some embodiments, the conventional architecture of a neural network processor (NPU) is as follows Figure 3 As shown, the control unit NCU is responsible for all instruction scheduling control of the NPU. Through efficient scheduling algorithms and control logic, it ensures that computing tasks are executed in an existing sequence to reduce waiting time. The data transmission module DMA realizes high-speed data transmission through the DMA controller. BUFFER: Responsible for storing feature maps, weights, and quantization parameters. By using high-speed on-chip storage, it reduces data transmission delays and improves access speed. The first computing module VECTOR: Responsible for vector operations and other operations besides convolution and matrix operations, such as common activation functions (ReLU, Sigmoid, Tanh, etc.), to improve computing efficiency. The second computing module PE ARRAY: A dedicated hardware unit to efficiently perform key core computing tasks such as matrix multiplication and convolution, supporting large-scale parallel computing.

[0027] However, in actual use, since the system can only process one neural network task at a time, some resources will inevitably not be fully utilized. Although the NPU itself has the ability to run in parallel, when each unit performs different tasks, there is a certain data dependency between the tasks. This dependency makes it impossible for the tasks to be executed completely in parallel. In this case, some units will inevitably be idle and unable to fully realize their computing potential. This data dependency is one of the main reasons for the limited utilization of NPU resources, because it hinders the complete parallel processing between tasks, thereby affecting the overall computing efficiency. Therefore, although the NPU has significant advantages in neural network calculations, in actual applications, the problem of insufficient resource utilization still needs to be solved to further improve system performance.

[0028] Therefore, the neural network processor (NPU) disclosed in the present invention is provided with a first control unit and a second control unit; the upgrade from a single control unit structure to a dual control unit design gives the NPU the ability to schedule two neural network instructions at the same time. This dual control unit architecture is designed to efficiently allocate and coordinate the use of computing resources, thereby significantly improving the work efficiency and resource utilization of the NPU. In the traditional NPU architecture, a single control core is used. Although the NPU has parallelism, due to certain dependencies between instructions, all functional modules cannot be completely executed in parallel. For example, when the NPU executes DMA and VECTOR instructions, the PE ARRAY module is in an idle state. However, if there are two control units, when the first control unit executes DMA and VECTOR instructions, the second control unit can execute PE ARRAY instructions. In this way, the NPU resources are fully utilized.

[0029] In the present disclosure, the software configures corresponding registers for the first control unit (NCU) and the second control unit (NCU) one by one through the bus system to achieve refined control and management of each control unit.

[0030] Specifically, it includes: configuring an initialization register to the first control unit, and configuring a boot register for startup; configuring an initialization register to the second control unit, and configuring a boot register for startup.

[0031] It should be noted that in this process, independent buffers are allocated to the two control units, that is, the first control unit and the second control unit are both provided with buffers; wherein, the capacity of the above-mentioned buffers is not fixed. Furthermore, the buffers of the above-mentioned first control unit include: a first weight buffer, a first quantization parameter buffer, and a first feature map buffer; a second weight buffer, a second quantization parameter buffer, and a second feature map buffer; wherein, the capacity of the first weight buffer is equal to the capacity of the second weight buffer; the capacity of the first quantization parameter buffer is equal to the capacity of the second quantization parameter buffer; and the capacity of the first feature map buffer is equal to or not equal to the capacity of the second feature map buffer.

[0032] Specifically, the buffer is divided into at least three parts, which respectively store images or feature maps, weights, and quantization parameters. The weight and quantization parameter buffers will be evenly distributed to two control units, and the specific capacity of the feature map buffer is determined according to the complexity and characteristics of the neural network model that each control unit needs to execute. This means that the size of the feature map buffer of the first control unit and the second control unit is not fixed, but according to the specific requirements of their tasks, the feature map buffer capacity of the two control units is dynamically adjusted by configuring the feature map storage offset address of the second control unit during the initialization phase. This flexible allocation strategy ensures that each core makes full use of memory resources when executing neural networks, thereby achieving better performance. In this way, the software not only optimizes the use of resources, but also improves the processing power and operating efficiency of the overall system.

[0033] The configuration of the initialization register to the second control unit stated above includes at least: configuring the feature map storage offset address in the offset address register of the second control unit; wherein the above-mentioned feature map storage offset address is used to determine the capacity of the feature map buffer allocated to the first control unit and the second control unit.

[0034] In some embodiments, the offset address is written to the offset address register of the second control unit via the AHB bus. If the first control unit executes the first neural network task, the first control unit is controlled to automatically execute the first neural network task, including:

[0035] The first step is to configure an enable register in the first control unit to enable the first control unit. The second step is to control the first control unit to retrieve a boot instruction from the boot register; wherein the boot instruction is used to load the relevant instructions of the first neural network. The third step is to automatically execute the first neural network task in response to the loading being completed.

[0036] In some embodiments, if the first control unit executes the second neural network task, controlling the first control unit to automatically execute the second neural network task includes:

[0037] The first step is to configure an enable register in the first control unit to enable the first control unit. The second step is to control the first control unit to retrieve a boot instruction from the boot register; the boot instruction is used to load the relevant instructions of the second neural network. The third step is to automatically execute the second neural network task in response to the completion of the loading.

[0038] As an example: the first neural network may be an image neural network; the second neural network may be a speech neural network.

[0039] If the second control unit executes the second neural network task, controlling the second control unit to automatically execute the second neural network task includes:

[0040] The first step is to configure an enable register for the second control unit to enable the second control unit. The second step is to control the second control unit to retrieve a boot instruction from the boot register; the boot instruction is used to load the relevant instructions of the second neural network. The third step is to automatically execute the second neural network task in response to the completion of the loading.

[0041] In some embodiments, if the second control unit executes the first neural network task, controlling the second control unit to automatically execute the first neural network task includes:

[0042] The first step is to configure an enable register in the second control unit to enable the second control unit. The second step is to control the second control unit to retrieve a boot instruction from the boot register; wherein the boot instruction is used to load the relevant instructions of the first neural network. The third step is to automatically execute the first neural network task in response to the completion of the loading.

[0043] Step S102: Control the first control unit to automatically execute the first neural network task.

[0044] In some embodiments, in terms of technical implementation, the first control unit executes a neural network of images for face recognition, gesture recognition, and posture estimation, to confirm user identity, understand user gesture commands, etc.

[0045] Step S103, controlling the second control unit to automatically execute a second neural network task; wherein the second neural network task is different from the first neural network task.

[0046] In some embodiments, the second control unit executes a speech neural network to understand the user's voice commands. The system makes the final control decision by integrating image and voice information. The second neural network task is different from the first neural network task.

[0047] It should be noted that this disclosure can be used in dual neural network application scenarios, particularly in high-demand areas such as image and voice networks and image and AI ISP networks. The technical solutions of this disclosure not only enable independent R&D and tape-out to achieve highly customized solutions, but also offer the option of licensing IP (intellectual property) for rapid integration into existing systems, significantly reducing time to market. In the image and voice network sector, this disclosure provides an efficient neural network architecture that significantly improves the accuracy and efficiency of image recognition and voice processing. For image and AI ISP networks, this disclosure combines advanced AI algorithms with image signal processing technology to significantly improve image quality while reducing power consumption and latency. Whether developing and tape-out in-house or purchasing IP licenses, this disclosure offers flexible options to meet the needs of diverse users. Users who develop and tape-out in-house can leverage the technical advantages of this disclosure to create competitive products; those who choose to purchase IP licenses can quickly integrate advanced technologies and enhance their products' market competitiveness. Furthermore, this disclosure supports expansion into a variety of application scenarios, such as autonomous driving, smart homes, and medical image analysis, offering users a wide range of application prospects. With the technical support of this disclosure, users can achieve technological breakthroughs in multiple fields and drive innovation and development in the industry.

[0048] In some optional implementations of some embodiments, the first control unit and the second control unit are further used to issue configuration requests to functional modules in the neural network processor; wherein the functional modules include a data transmission module, a first computing module, and a second computing module; the data transmission module, the first computing module, and the second computing module are all configured with corresponding arbitrators; the above-mentioned arbitrators are used to coordinate the configuration requests of the first control unit and the second control unit.

[0049] It should be noted that the above-mentioned arbitrator can coordinate the configuration requests of the first control unit and the second control unit based on the time rotation mechanism, and can also adopt other arbitration mechanisms to coordinate the configuration requests of the first control unit and the second control unit, which is not specifically limited here.

[0050] In some embodiments, as Figure 3 As shown in the figure: In the dual-control unit core parallel processing architecture, the two control units independently execute their own instructions. During execution, each control unit may issue specific configuration requests to functional modules outside the control unit based on task requirements. These functional modules outside the core include three key modules: data transmission module (DMA module), first calculation module (VECTOR module) and second calculation module (PE ARRAY module). These modules are crucial components in neural network calculations, responsible for tasks such as data transmission, vector operations, convolution and matrix calculations.

[0051] Since two control units may issue configuration requests to the same functional module at the same time or at similar time points (for example, they both need to access DMA for data transfer), and these functional modules can only process one request at a time, an arbitration mechanism is required to ensure that the requests are processed in order.

[0052] To address this issue, the NPU incorporates an arbiter, responsible for coordinating configuration requests from two control units. When two control units simultaneously issue configuration requests for the same functional module, the arbiter determines which control unit's request should be processed first based on a time-based, round-robin mechanism. The arbiter immediately passes the selected control unit's configuration request to the corresponding functional module (such as DMA, VECTOR, or PE ARRAY) to ensure timely execution.

[0053] The arbiter also features an internal cache to temporarily store unexecuted core configuration requests. While one core's request is being processed, another core's configuration request is temporarily stored in the arbiter's cache. After the current request is processed, the cached request is sent to the functional module for execution. This ensures that every core's request is processed, preventing configuration information from being lost due to resource conflicts.

[0054] Through this arbitration mechanism, the system can efficiently coordinate access to limited functional modules by multiple control units, avoiding resource conflicts and ensuring that each control unit's configuration request is processed at the appropriate time. This design not only enhances the system's parallel processing capabilities but also significantly reduces delays caused by resource competition, thereby improving overall computing performance and smoother task execution.

[0055] As an example: the first control unit sends a request to the arbitrator first, and the arbitrator will send the request of the first control unit. Later, if the second control unit also sends a request to the arbitrator, but the request of the first control unit to the corresponding functional modules (data transmission module (DMA module), first calculation module (VECTOR module) and second calculation module (PE ARRAY module)) has not yet completed its work. Then the request of the second control unit will be temporarily stored in the cache by the arbitrator. After the request sent by the first control unit is completed. The arbitrator will take out the request of the second control unit from the cache and send it to the corresponding functional module.

[0056] In summary, when the two control units execute instructions independently, the arbitrator arbitrates and schedules configuration requests for the same functional module, ensuring efficient use of system resources.

[0057] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.

[0058] The following are embodiments of the apparatus disclosed herein, which can be used to implement the method embodiments disclosed herein. For details not disclosed in the apparatus embodiments disclosed herein, please refer to the method embodiments disclosed herein.

[0059] Figure 4 4 is a schematic diagram of the structure of some embodiments of a neural network processor according to the present disclosure, wherein the neural network processor includes: a first control unit 401, a second control unit 402, a functional module 403 and an arbitrator 404; wherein:

[0060] The first control unit 401 is used to execute a first neural network task;

[0061] The second control unit 402 is used to execute a second neural network task;

[0062] The functional module 403 is used for data transmission and calculation;

[0063] The arbitrator 404 is used to coordinate the configuration requests of the first control unit 401 and the second control unit 402 .

[0064] The NPU in this embodiment can be applied to a reconfigurable computing chip. The reconfigurable computing chip includes the NPU. In this embodiment, the NPU includes a dual-core NCU, namely, NCU Core 1 and NCU Core 2. The on-chip software of the reconfigurable computing chip configures the NCU.

[0065] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A task processing method based on dual control units, characterized in that: The method comprises: Configuring initialization registers for the first control unit and the second control unit, respectively; wherein the first control unit and the second control unit are control units in the same neural network processor; configuring the initialization register for the second control unit comprises at least: configuring a feature map storage offset address in an offset address register of the second control unit; wherein the feature map storage offset address is used to determine the capacity of the feature map buffer allocated to the first control unit and the second control unit; controlling the first control unit to automatically execute the first neural network task; Control the second control unit to automatically execute a second neural network task; wherein the second neural network task is different from the first neural network task.

2. The task processing method based on dual control units according to claim 1, characterized in that: Configuring initialization registers for the first control unit and the second control unit respectively includes: Configuring an initialization register for the first control unit and configuring a boot register for startup; The initialization register is configured in the second control unit, and the boot register is configured for startup.

3. The task processing method based on dual control units according to claim 1, characterized in that: If the first control unit executes the first neural network task, controlling the first control unit to automatically execute the first neural network task includes: Configuring an enable register for the first control unit to enable the first control unit; Controlling the first control unit to retrieve a boot instruction from the boot register; wherein the boot instruction is used to load relevant instructions of the first neural network; In response to the completion of loading, the first neural network task is automatically executed.

4. The task processing method based on dual control units according to claim 1, characterized in that: If the second control unit executes the second neural network task, controlling the second control unit to automatically execute the second neural network task includes: Configuring an enable register for the second control unit to enable the second control unit; Controlling the second control unit to retrieve a boot instruction from the boot register; wherein the boot instruction is used to load relevant instructions of the second neural network; In response to the completion of loading, the second neural network task is automatically executed.

5. The task processing method based on dual control units according to claim 1, characterized in that: The first control unit and the second control unit are both provided with a buffer zone; wherein the capacity of the buffer zone is not fixed.

6. The task processing method based on dual control units according to claim 5, characterized in that: The buffer of the first control unit includes: a first weight buffer, a first quantization parameter buffer and a first feature map buffer; a second weight buffer, a second quantization parameter buffer and a second feature map buffer; wherein the capacity of the first weight buffer is equal to the capacity of the second weight buffer; the capacity of the first quantization parameter buffer is equal to the capacity of the second quantization parameter buffer; the capacity of the first feature map buffer is equal to or not equal to the capacity of the second feature map buffer.

7. The task processing method based on dual control units according to claim 1, characterized in that: The first control unit and the second control unit are further configured to issue configuration requests to functional modules in the neural network processor; wherein the functional modules include a data transmission module, a first computing module, and a second computing module; the data transmission module, the first computing module, and the second computing module are each configured with a corresponding arbitrator; the arbitrator is configured to coordinate the configuration requests of the first control unit and the second control unit.

8. A neural network processor for implementing the method according to any one of claims 1 to 7, characterized in that: The neural network processor includes: a first control unit, a second control unit, a functional module and an arbitrator; wherein: The first control unit is used to perform a first neural network task; The second control unit is used to perform a second neural network task; The functional modules are used for data transmission and calculation; The arbitrator is used to coordinate configuration requests of the first control unit and the second control unit.

9. A computer system, characterized in that: The computer system includes the neural network processor of claim 8.

10. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 7 when executing the computer program.