Neural network processor monitoring system, method and device and electronic equipment

By adding a detection module to each target computing unit in the computing cluster module of the neural network processor, the operating status is monitored and fed back in real time, which solves the problem of long delay in the CPU being informed of NPU abnormalities and realizes efficient status monitoring.

CN120704972APending Publication Date: 2025-09-26SHANGHAI LIXIANG AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410331823.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the prior art, the central processing unit (CPU) has a long delay in learning that the neural network processor (NPU) is operating in an abnormal state, making it impossible to monitor in real time.

Method used

A detection module is added to each target computing unit in the computing cluster module of the neural network processor. The operating status information is obtained through the detection module and sent to the monitoring module. The monitoring module triggers a central processing unit interrupt when an abnormality is detected.

Benefits of technology

It realizes real-time feedback of the operating status of the neural network processor, reduces the delay of the central processor in being informed of abnormalities, and improves monitoring efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704972A_ABST
    Figure CN120704972A_ABST
Patent Text Reader

Abstract

The invention discloses a neural network processor monitoring system, method and device and electronic equipment, and relates to the technical field of equipment, a detection module is added to each target computing unit included in a computing cluster module of a neural network processor, and each detection module is connected with a monitoring module; when the target calculation unit executes the model algorithm task, a first target operation state of the target calculation unit is collected based on the detection module, the first target operation state is sent to the monitoring module, and the monitoring module receives the first target operation state and sends the first target operation state to the target calculation unit; and when the monitoring module determines that the first target operation state is abnormal, the central processing unit is triggered to be interrupted, so that the monitoring module can feed back the operation state of the neural network processor to the central processing unit in real time without waiting for the completion of the execution of the target instruction. And the time delay when the central processing unit knows that the operation state of the neural network processor is abnormal is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of equipment technology, and in particular to a neural network processor monitoring system, method, device and electronic equipment. Background Art

[0002] Currently, the execution of model algorithm tasks is mainly achieved through the central processing unit (CPU), which pushes a large number of different types of model algorithm tasks to the neural processing unit (NPU) so that the NPU can perform inference and operation on the model algorithm tasks. When various model algorithm tasks are running on the NPU, it is extremely important for the CPU to monitor the operating status of the NPU.

[0003] In the related technology, when the CPU monitors the running status of the NPU, the CPU is used to dispatch different types of model algorithm tasks to the host interface block processor (HCPU). After receiving the model algorithm task, the HCPU sends a Tile instruction to the Tile module so that the computing unit in the Tile module can calculate the model algorithm task. The Tile module contains all the computing units of the NPU but does not contain a monitoring unit. Therefore, the monitoring of the running status of the NPU can only be achieved through the HCPU, that is, the HCPU monitors the running status of the NPU in executing various model algorithm tasks. After the HCPU detects that the running status of the NPU in executing various model algorithm tasks is abnormal, it notifies the CPU and the CPU takes corresponding measures to realize the CPU monitoring the running status of the NPU in executing various model algorithm tasks. However, since the scheduling of model algorithm tasks by the HCPU is an atomic operation, it is necessary to complete the execution of all Tile instructions before the abnormal running status of the NPU can be fed back to the CPU, which in turn causes a long delay in the CPU learning that the running status of the NPU is abnormal. Summary of the Invention

[0004] The present disclosure provides a neural network processor monitoring system, method, device, and electronic device, the main purpose of which is to solve the problem of long delay in the CPU learning the abnormal operating status of the NPU in the prior art.

[0005] According to a first aspect of the present disclosure, a neural network processor monitoring system is provided, comprising:

[0006] A computing cluster module includes multiple target computing units, each of which includes a detection module. The computing cluster module is used to receive target instructions, and the target instructions are used to drive the computing cluster module to use the target computing unit to execute the model algorithm task pushed by the central processing unit to the neural network processor;

[0007] The detection module is configured to obtain first target operating status information of the corresponding target computing unit and send the first target operating status information to the monitoring module;

[0008] The monitoring module connected to each detection module is used to receive the first target operation status information and trigger the central processing unit to interrupt when determining that the first target operation status information is abnormal.

[0009] Optionally, the central processing unit further includes:

[0010] A control state host interface block processor connected to the central processing unit and the host interface block processor respectively is used to monitor and obtain the second target operating state information of the host interface block processor, and trigger the central processing unit interrupt when it is determined that the second target operating state information is abnormal, wherein the host interface block processor is a structural component of the neural network processor, and the host interface block processor is used to schedule the model algorithm task, and generate corresponding target instructions according to the model algorithm task, and send the target instructions to the computing cluster module.

[0011] Optionally, the detection module is a monitoring IP block of the target computing unit, and the monitoring IP block is an integrated circuit design component for monitoring and managing the target computing unit; the monitoring module is connected to the central processing unit via an on-chip bus.

[0012] According to a second aspect of the present disclosure, a method for monitoring a neural network processor is provided, comprising:

[0013] Receive model algorithm tasks;

[0014] During the execution of the model algorithm task, obtaining first target operating state information;

[0015] The first target operating state information is sent to a monitoring module, so that the monitoring module sends a first interrupt signal to a central processing unit when determining that the first target operating state information is abnormal.

[0016] Optionally, after receiving the model algorithm task, the method includes:

[0017] Acquire second target running status information, where the second target running status information is running status information of a host interface block processor;

[0018] The second target operating state information is sent to the control state host interface block processor, so that the control state host interface block processor sends a second interrupt signal to the central processing unit when determining that the second target operating state information is abnormal.

[0019] According to a third aspect of the present disclosure, a method for monitoring a neural network processor is provided, comprising:

[0020] Sending the model algorithm task to the neural network processor, so that the neural network processor obtains first target operating state information during the execution of the model algorithm task;

[0021] In response to receiving a first interrupt signal sent by a monitoring module, an interrupt operation is performed, wherein the monitoring module is used to trigger the first interrupt signal when it is determined that the first target operating status information sent by the neural network processor is abnormal.

[0022] Optionally, after sending the model algorithm task to the neural network processor, the method includes:

[0023] In response to receiving a second interrupt signal sent by the control state host interface block processor, an interrupt operation is performed, and the control state host interface block processor is used to trigger the second interrupt signal when it is determined that there is an abnormality in the second target operating status information sent by the neural network processor, and the second target operating status information is the operating status information of the host interface block processor.

[0024] Optionally, after the interrupt operation is performed, the method includes:

[0025] collecting the first target operating state information received by the monitoring module and / or the second target operating state information received by the control state host interface block processor;

[0026] The first target running status information and / or the second target running status information are saved to a target server.

[0027] Optionally, before sending the model algorithm task to the neural network processor, the method includes:

[0028] Determining whether the computing resources of the neural network processor are sufficient;

[0029] If the computing resources of the neural network processor are sufficient, controlling the host interface block processor to run a monitoring algorithm, wherein the monitoring algorithm is used to predict whether an abnormality will occur in the operating state of the neural network processor;

[0030] If the computing resources of the neural network processor are insufficient, the central processing unit is controlled to run the monitoring algorithm.

[0031] According to a fourth aspect of the present disclosure, there is provided an apparatus for monitoring a neural network processor, comprising:

[0032] A receiving unit, used for receiving model algorithm tasks;

[0033] A first acquisition unit is used to acquire first target operating state information during the execution of the model algorithm task;

[0034] The first sending unit is configured to send the first target operating status information to the monitoring module, so that the monitoring module sends a first interrupt signal to the central processing unit when determining that the first target operating status information is abnormal.

[0035] Optionally, the device further includes:

[0036] A second acquiring unit, configured to acquire second target running status information, where the second target running status information is running status information of a host interface block processor;

[0037] The second sending unit is configured to send the second target operating state information to the control state host interface block processor, so that the control state host interface block processor sends a second interrupt signal to the central processing unit when determining that the second target operating state information is abnormal.

[0038] According to a fifth aspect of the present disclosure, there is provided an apparatus for monitoring a neural network processor, comprising:

[0039] a sending unit, configured to send the model algorithm task to the neural network processor, so that the neural network processor obtains first target operating state information during the execution of the model algorithm task;

[0040] An execution unit is used to perform an interrupt operation in response to receiving a first interrupt signal sent by a monitoring module, wherein the monitoring module is used to trigger the first interrupt signal when determining that the first target operating status information sent by the neural network processor is abnormal.

[0041] Optionally, the execution unit is further configured to:

[0042] In response to receiving a second interrupt signal sent by the control state host interface block processor, an interrupt operation is performed, and the control state host interface block processor is used to trigger the second interrupt signal when it is determined that there is an abnormality in the second target operating status information sent by the neural network processor, and the second target operating status information is the operating status information of the host interface block processor.

[0043] Optionally, the device includes:

[0044] a collection unit, configured to collect the first target operating state information received by the monitoring module and / or the second target operating state information received by the control state host interface block processor;

[0045] A saving unit is configured to save the first target running status information and / or the second target running status information to a target server.

[0046] Optionally, the device includes:

[0047] a determination unit, configured to determine whether the computing resources of the neural network processor are sufficient;

[0048] a control unit, configured to control the host interface block processor to execute a monitoring algorithm when the computing resources of the neural network processor are sufficient, wherein the monitoring algorithm is configured to predict whether an abnormality will occur in the operating state of the neural network processor;

[0049] The control unit is further configured to control the central processing unit to execute the monitoring algorithm when the computing resources of the neural network processor are insufficient. According to a sixth aspect of the present disclosure, an electronic device is provided, comprising:

[0050] at least one processor; and

[0051] a memory communicatively connected to the at least one processor; wherein,

[0052] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the second aspect or the third aspect.

[0053] According to a seventh aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the second or third aspect.

[0054] According to an eighth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the second or third aspect above.

[0055] The present disclosure provides a neural network processor monitoring system, method, device and electronic device, a computing cluster module, comprising multiple target computing units, each target computing unit including a detection module, the computing cluster module being used to receive target instructions, the target instructions being used to drive the computing cluster module to use the target computing unit to execute the model algorithm task pushed by the central processing unit to the neural network processor; the detection module being used to obtain the first target operating status information of the corresponding target computing unit and to send the first target operating status information to the monitoring module; the monitoring module connected to each detection module being used to receive the first target operating status information and, when determining that the first target operating status information is abnormal, triggering an interrupt to the central processing unit. Compared with the related art, by adding a detection module to each target computing unit included in the computing cluster module of the neural network processor, and connecting each detection module with the monitoring module, when the target computing unit executes the model algorithm task, the first target operating state of the target computing unit is collected based on the detection module, and the first target operating state is sent to the monitoring module. The monitoring module receives the first target operating state, and when the monitoring module determines that the first target operating state is abnormal, it triggers a central processing unit interrupt, so that the monitoring module can feedback the operating state of the neural network processor to the central processing unit in real time without waiting for the target instruction to be executed, thereby reducing the delay for the central processing unit to know that the operating state of the neural network processor is abnormal.

[0056] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0058] Figure 1 A schematic diagram of the structure of a neural network processor monitoring system provided by an embodiment of the present disclosure;

[0059] Figure 2 A schematic diagram of the structure of another neural network processor monitoring system provided by an embodiment of the present disclosure;

[0060] Figure 3 A flowchart of a method for monitoring a neural network processor provided by an embodiment of the present disclosure;

[0061] Figure 4 A flowchart of another method for monitoring a neural network processor provided by an embodiment of the present disclosure;

[0062] Figure 5 A schematic diagram of a program architecture for executing a neural network processor monitoring method provided by an embodiment of the present disclosure;

[0063] Figure 6 A schematic diagram of the structure of a device for monitoring a neural network processor provided in an embodiment of the present disclosure;

[0064] Figure 7 A schematic structural diagram of another apparatus for monitoring a neural network processor provided by an embodiment of the present disclosure;

[0065] Figure 8 A schematic structural diagram of another apparatus for monitoring a neural network processor provided by an embodiment of the present disclosure;

[0066] Figure 9 A schematic structural diagram of another apparatus for monitoring a neural network processor provided by an embodiment of the present disclosure;

[0067] Figure 10 A schematic block diagram of an exemplary electronic device 600 is provided for an embodiment of the present disclosure. DETAILED DESCRIPTION

[0068] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0069] The following describes the neural network processor monitoring system, method, device and electronic device according to the embodiments of the present disclosure with reference to the accompanying drawings.

[0070] Figure 1 A schematic diagram of the structure of a neural network processor monitoring system provided in an embodiment of the present disclosure.

[0071] like Figure 1 As shown, the system includes: a computing cluster module 121, a detection module 1211, a monitoring module 11, a central processing unit 13, and a neural network processor 12;

[0072] The computing cluster module 121 includes multiple target computing units, each of which includes a detection module 1211. The computing cluster module 121 is used to receive target instructions, and the target instructions are used to drive the computing cluster module 121 to use the target computing unit to execute the model algorithm task pushed by the central processing unit 13 to the neural network processor 12;

[0073] The detection module 1211 is configured to obtain first target operating status information of the corresponding target computing unit and send the first target operating status information to the monitoring module 11;

[0074] The monitoring module 11 connected to each detection module 1211 is used to receive the first target operating status information and trigger the central processing unit 13 to interrupt when determining that the first target operating status information is abnormal.

[0075] As a refinement of the above embodiment, please refer to Figure 1 , Figure 1 The connection relationship between the central processing unit 13 and the neural network processor 12 is shown. Before the monitoring module 11 monitors the neural network processor 12, the central processing unit 13 distributes the model algorithm task to the neural network processor 12. After receiving the model algorithm task, the neural network processor 12 generates a target instruction corresponding to the model algorithm task and sends the target instruction to the computing cluster module 121, so that the computing cluster module 121 uses the target computing unit to execute the model algorithm task. Figure 1 It can be seen that the computing cluster module 121 is a component of the neural network processor 12. In addition, the computing cluster module 121 includes all computing units of the neural network processor 12. In order to monitor the operating status of the neural network processor 12 during the execution of the model algorithm task by the target computing unit, it is necessary to monitor the first target operating status information of the target computing unit included in the neural network processor 12. In order to achieve the above purpose, each target computing unit includes a detection module 1211, so as to obtain the first target operating status information of the corresponding target computing unit through the detection module 1211, and send the first target operating status information to the monitoring module 11. The monitoring module 11 is connected to each detection module 1211. When the monitoring module 11 determines that the first target operating status information is abnormal, it triggers the interruption of the central processing unit 13, thereby realizing that when the operating status of the neural network processor 12 is abnormal, the central processing unit 13 is notified, and the central processing unit 13 monitors the operating status of the neural network processor 12.

[0076] The neural network processor monitoring system provided by the present disclosure includes: a computing cluster module 121, which includes multiple target computing units, each target computing unit includes a detection module 1211, the computing cluster module 121 is used to receive target instructions, and the target instructions are used to drive the computing cluster module 121 to use the target computing unit to execute the model algorithm task pushed to the neural network processor 12 by the central processing unit 13; the detection module 1211 is used to obtain the first target operating status information of the corresponding target computing unit and send the first target operating status information to the monitoring module 11; the monitoring module 11 connected to each detection module 1211 is used to receive the first target operating status information and trigger the central processing unit 13 to interrupt when determining that the first target operating status information is abnormal. Compared with the related art, by adding a detection module 1211 to each target computing unit included in the computing cluster module 121 of the neural network processor 12, and connecting each detection module 1211 with the monitoring module 11, when the target computing unit executes the model algorithm task, the first target operating state of the target computing unit is collected based on the detection module 1211, and the first target operating state is sent to the monitoring module 11. The monitoring module 11 receives the first target operating state. When the monitoring module 11 determines that the first target operating state is abnormal, it triggers the central processing unit 13 to interrupt, so that the monitoring module 11 can feedback the operating state of the neural network processor 12 to the central processing unit 13 in real time without waiting for the target instruction to be executed, thereby reducing the delay for the central processing unit 13 to know that the operating state of the neural network processor 12 is abnormal.

[0077] Figure 2 This is a structural diagram of another neural network processor monitoring system provided by an embodiment of the present disclosure, such as Figure 2 As shown, Figure 1 The system shown is compared Figure 2 In the system shown, the central processor 13 also includes:

[0078] The control state host interface block processor 131 connected to the central processing unit and the host interface block processor respectively is used to monitor and obtain the second target operating state information of the host interface block processor 122, and trigger the central processing unit 13 to interrupt when it is determined that the second target operating state information is abnormal, wherein the host interface block processor 122 is a structural component of the neural network processor 12, and the host interface block processor 122 is used to schedule the model algorithm task, and generate corresponding target instructions according to the model algorithm task, and send the target instructions to the computing cluster module 121.

[0079] As a refinement of the above embodiment, the control state host interface block processor 131 is added to the central processing unit 13, so that the task of monitoring the host interface block processor 122 of the neural network processor 12 is undertaken by the control state host interface block processor 131, thereby reducing the load of the central processing unit 13 and realizing the monitoring of the operating status of the host interface block processor 122. The method is the same as the method in which the aforementioned monitoring module 11 triggers the interruption of the central processing unit 13. When the control state host interface block processor 131 determines that the second target operating status information is abnormal, the central processing unit 13 is triggered to interrupt.

[0080] Depend on Figure 2 It can be seen that the host interface block processor 122 is a structural component of the neural network processor 12. The host interface block processor 122 is used to schedule the model algorithm task distributed by the central processor 13 to the neural network processor 12. The host interface block processor 122 generates corresponding target instructions according to the model algorithm task, and sends the target instructions to the computing cluster module 121, so that the target instructions drive the computing cluster module 121 to use the model algorithm task of the target computing unit.

[0081] In some embodiments, the monitoring module 11 includes a state machine inside, and the monitoring module 11 is also used to debug and trace the host interface block processor 122. The monitoring module 11 includes a JTAG interface, and the debug and trace functions can greatly improve development efficiency, decouple the central processing unit 13, and reduce the interference of the central processing unit 13 on the neural network processor 12.

[0082] To facilitate understanding of the above embodiment, this embodiment further explains the above embodiment. In order to debug and trace the host interface block processor 122 through the monitoring module 11, the monitoring module 11 is embedded with a JTAG interface. By connecting the debugging peripheral to the JTAG interface of the monitoring module 11, direct debugging and monitoring of the host interface block processor 122 can be achieved. Among them, debug is the function of diagnosing system faults, and trace is the function of recording the system operation trajectory.

[0083] In some embodiments, the debug and trace control signals in the monitoring module 11 may be shielded from the JTAG interface and may only be controlled by the central processing unit 13 during the mass production stage.

[0084] As a refinement of the above embodiment, the detection module 1211 is a monitoring IP block of the target computing unit, and the monitoring IP block is an integrated circuit design component for monitoring and managing the target computing unit; the monitoring module 11 is connected to the central processing unit 13 through an on-chip bus.

[0085] As a refinement of the above embodiment, the monitoring IP block refers to an integrated circuit module with specific functions. Furthermore, the monitoring IP block refers to an integrated circuit design component used to monitor and manage the target computing unit, which can be embedded in a chip to monitor information such as the chip's operating status, performance, temperature, and power consumption. It should be understood that the above description of the monitoring IP block is merely exemplary and does not constitute a limitation of the present disclosure. For example, the monitoring module 11 can be connected to the central processing unit 13 via an AXI protocol bus to achieve communication between the monitoring module 11 and the central processing unit 13.

[0086] Corresponding to the aforementioned neural network processor monitoring system, the present invention also provides a method for monitoring a neural network processor. Since the method embodiment of the present invention corresponds to the aforementioned system embodiment, details not disclosed in the method embodiment can be referred to the aforementioned system embodiment and will not be further described in this invention.

[0087] Figure 3 This is a flow chart of a method for monitoring a neural network processor provided by an embodiment of the present disclosure. The method is applied to a neural network processor in the neural network processor monitoring system. The specific steps are as follows:

[0088] Step 201, receiving a model algorithm task;

[0089] As a refinement of the above step 201, the central processing unit sends the model algorithm task to the neural network processor, and the neural network processor receives the model algorithm task and executes the model algorithm task.

[0090] Step 202: obtaining first target operation status information during execution of the model algorithm task;

[0091] As a refinement of the above-mentioned step 202, in the process of executing the model algorithm task, the detection module included in each target computing unit in the computing cluster module in the neural network processor will obtain the operating status of its corresponding target computing unit, thereby obtaining the first target operating status information, which is the operating status information of each target computing unit.

[0092] Step 203: Send the first target operating status information to a monitoring module, so that the monitoring module sends a first interrupt signal to a central processing unit when determining that the first target operating status information is abnormal.

[0093] As a refinement of step 203, in order to reduce the load on the central processing unit, the monitoring module determines whether the first target operating status information is abnormal. When the monitoring module determines that the first target operating status information is abnormal, the monitoring module generates the first interrupt signal and sends the first interrupt signal to the central processing unit. The central processing unit responds to the first interrupt signal and performs corresponding processing based on the abnormal first target operating status information.

[0094] The method for monitoring a neural network processor provided by the present disclosure adds a detection module to each target computing unit included in the computing cluster module of the neural network processor, and connects each detection module with the monitoring module. When the target computing unit executes the model algorithm task, the first target operating state of the target computing unit is collected based on the detection module, and the first target operating state is sent to the monitoring module. The monitoring module receives the first target operating state. When the monitoring module determines that the first target operating state is abnormal, the monitoring module triggers a central processing unit interrupt, so that the monitoring module can feedback the operating state of the neural network processor to the central processing unit in real time without waiting for the target instruction to be executed, thereby reducing the delay for the central processing unit to know that the operating state of the neural network processor is abnormal.

[0095] As a refinement of the embodiment of the present disclosure, after receiving the model algorithm task, the method can also adopt but is not limited to the following implementation methods, for example: obtaining second target operating status information, the second target operating status information is the operating status information of the host interface block processor; sending the second target operating status information to the control state host interface block processor, so that the control state host interface block processor sends a second interrupt signal to the central processing unit when determining that there is an abnormality in the second target operating status information.

[0096] In order to monitor the host interface block processor in the neural network processor, the operating status information of the host interface block processor, that is, the second target operating status information, is sent to the control state host interface block processor, and the control state host interface block processor determines whether there is an abnormality in the second target operating status information. If it is determined that there is an abnormality in the second target operating status information, the control state host interface block processor sends the second interrupt signal to the central processing unit to trigger the central processing unit interrupt.

[0097] Figure 4 This is a flow chart of another method for monitoring a neural network processor provided by an embodiment of the present disclosure. The method is applied to a central processor in a neural network processor monitoring system. The specific steps are as follows:

[0098] Step 301: Sending a model algorithm task to a neural network processor, so that the neural network processor obtains first target operating state information during the execution of the model algorithm task;

[0099] As a refinement of the above step 301, the model algorithm task is sent to the neural network processor so that the neural network processor executes the model algorithm task process and obtains the first target operating status information during the execution process.

[0100] Step 302: Execute an interrupt operation in response to receiving a first interrupt signal sent by a monitoring module, wherein the monitoring module is configured to trigger the first interrupt signal when determining that the first target operating status information sent by the neural network processor is abnormal.

[0101] As a refinement of step 302, the central processing unit executes an interrupt operation in response to the first interrupt signal, the interrupt operation including but not limited to interrupting the transmission of the model algorithm task. The first interrupt signal is an interrupt signal sent by the monitoring module when it determines that the first target operating state information sent by the neural network processor is abnormal.

[0102] In some embodiments, in order for the central processing unit to respond to the first interrupt signal of the monitoring module and perform corresponding processing, the central processing unit needs to register an interrupt function of a type corresponding to the monitoring module.

[0103] As a refinement of an embodiment of the present disclosure, after sending the model algorithm task to the neural network processor, the method may also adopt but is not limited to the following implementation methods, for example: in response to receiving a second interrupt signal sent by the control state host interface block processor, executing an interrupt operation, the control state host interface block processor is used to trigger the second interrupt signal when determining that there is an abnormality in the second target operating status information sent by the neural network processor, the second target operating status information is the operating status information of the host interface block processor.

[0104] In order to facilitate understanding of the above embodiment, this embodiment describes the above embodiment in detail, including: after receiving the second interrupt signal, the central processing unit performs an interrupt operation, which is the same as the receiving process of the first interrupt signal in the above embodiment, and this embodiment will not elaborate on this. The second interrupt signal is sent by the control state host interface block processor when it determines that there is an abnormality in the second target operating state information sent by the neural network processor.

[0105] In some embodiments, in order for the central processor to respond to the second interrupt signal of the control state host interface block processor and perform corresponding processing, the central processor needs to register an interrupt function of a type corresponding to the control state host interface block processor.

[0106] As a refinement of the above embodiment, after the execution interrupt operation, the method can also adopt but is not limited to the following implementation methods, for example: collecting the first target operating status information received by the monitoring module and / or the second target operating status information received by the control status host interface block processor; saving the first target operating status information and / or the second target operating status information to the target server.

[0107] In some embodiments, after the monitoring module sends the interrupt signal to the central processing unit, the central processing unit collects the first target operating status information from the monitoring module and saves the first target operating status information to the target server.

[0108] In some embodiments, after the monitoring module sends the interrupt signal to the central processing unit, if the control state host interface block processor also generates an interrupt signal and sends it to the central processing unit, the first target operating status information and the second target operating status information are collected, and the first target operating status information and the second target operating status information are saved to the target server.

[0109] In some embodiments, if only the control state host interface block processor generates an interrupt signal to send to the central processing unit, the second target operating state information is collected and saved to the target server.

[0110] As a refinement of the above embodiment, before sending the model algorithm task to the neural network processor, the method can also adopt but is not limited to the following implementation methods, for example: judging whether the computing resources of the neural network processor are sufficient; if the computing resources of the neural network processor are sufficient, running a monitoring algorithm based on the host interface block processor, and the monitoring algorithm is used to predict whether the operating status of the neural network processor will be abnormal; if the computing resources of the neural network processor are insufficient, controlling the central processing unit to run the monitoring algorithm.

[0111] In some embodiments, the central processor determines whether the computing resources of the neural network processor are sufficient. If the computing resources of the neural network processor are sufficient, the host interface block processor runs a monitoring algorithm, which is used to predict whether the operating status of the neural network processor will be abnormal. The monitoring algorithm is a neural network model.

[0112] As a refinement of the above embodiment, after determining whether the computing resources of the neural network processor are sufficient, the method can also adopt but is not limited to the following implementation methods, for example: if the computing resources of the neural network processor are insufficient, the monitoring algorithm is run based on the central processing unit.

[0113] In some embodiments, if the computing resources of the neural network processor are insufficient, the monitoring algorithm is run based on the central processing unit.

[0114] As a refinement of the above embodiment, after saving the first target operating status information and / or the second target operating status information to the target server, the method can also adopt but is not limited to the following implementation methods, for example: training the monitoring algorithm based on the first target operating status information and / or the second target operating status information to obtain a trained monitoring algorithm.

[0115] The monitoring algorithm is continuously trained using the saved first target operating state information and / or the saved second target operating state information to obtain a trained monitoring algorithm, thereby continuously improving the accuracy of the prediction results of the monitoring algorithm.

[0116] In some embodiments, when the monitoring algorithm predicts that the NPU may generate an exception, it immediately notifies the CPU. The CPU then processes the exception information based on the monitoring model's reasoning feedback. Exceptions include, but are not limited to, hardware failures (power supply, temperature issues), data anomalies, performance anomalies, memory anomalies, and power consumption anomalies. If the NPU does not encounter an exception, it executes the inference tasks assigned by the CPU in sequence.

[0117] In some embodiments, after completing the execution of the model algorithm task, the NPU returns the execution result of the model algorithm task to the algorithm application.

[0118] Figure 5 A schematic diagram of a program architecture for executing a neural network processor monitoring method provided by an embodiment of the present disclosure is provided. Figure 5 The figure shows that the program modules running on the CPU include the NPU Driver, which is the NPU driver, and the AIDriver, which is a program module specifically used to process artificial intelligence tasks. This module can run on the CPU to perform various AI-related tasks, such as monitoring algorithms. Linux is the operating system running on the CPU, and the program modules running on the HCPU include the Task Queue, a real-time operating system (RTOS), and the program modules running on the computing cluster module, including Ops, or model operators. Task represents the distribution of tasks to the NPU, dispatch refers to assigning tasks or requests to appropriate processes, threads, or services for processing, and watch refers to predicting the operating status of the NPU when executing model algorithm tasks.

[0119] In some embodiments, the computing cluster module is represented as a Tile module, the target instruction is represented as a Tile instruction, and the detection module is represented as a watch module. The method for monitoring a neural network model further includes: creating two kernel threads, namely a first target kernel thread and a second target kernel thread, and CPU registering Monitor and CSH type interrupt functions during NPU driver initialization. The NPU driver initialization process includes establishing a communication connection with the HCPU, powering on the NPU, negotiating shared memory, etc. The first target kernel thread and the second target kernel thread are respectively suspended, and the NPU driver main thread waits for a model algorithm task, which is an inference task that requires the NPU to execute. After the NPU driver main thread receives the model algorithm task, it wakes up the first target kernel thread and the second target kernel thread, and distributes the inference task to the HCPU. The HCPU sends a Ti le instruction according to the inference task to enable the Ti le module to execute the model algorithm task; after the first target kernel thread is awakened, it collects the status data of the NPU runtime from the CSH module, i.e., the control state host interface block processor, and the Monitor module, i.e., the monitoring module. The status data referred to here includes the first target running status information and the second target running status information; after the second target kernel thread is awakened, it first determines the load of the current NPU resources. If the resources are sufficient, the monitoring algorithm is sent to the idle HCPU for execution; if the resources are insufficient, the monitoring algorithm is executed on the thread of the current CPU. The monitoring algorithm will predict the abnormal behavior of the NPU based on the status information of the NPU runtime; when it is predicted that the NPU may send an exception, the CPU is immediately notified, and the CPU performs corresponding processing based on the abnormal information feedback from the monitoring model inference. The abnormal information includes: hardware failure (power supply, temperature problem), data abnormality, performance abnormality, memory abnormality, power consumption abnormality, etc.; if the monitoring algorithm does not predict the abnormality, the NPU will execute the model algorithm tasks distributed by the CPU in sequence; if the monitoring algorithm does not predict the abnormality but the CSH module or the Monitor module has triggered the CPU interrupt, the interrupt function will record the current status information and then notify the first target kernel thread. The first target kernel thread will collect the data during the NPU runtime. When the NPU model algorithm task is completed, the first target kernel thread will return the collected NPU data to the collection application, and the collection application will package the data and upload it to the cloud. The first target kernel thread and the second target kernel thread will suspend after the inference task is completed. After the model algorithm task is calculated, the NPU drives the main thread to obtain data and return the data to the algorithm application.

[0120] In summary, this embodiment can achieve the following effects:

[0121] 1. The monitoring module can provide real-time feedback on the operating status of the neural network processor to the central processing unit without waiting for the target instruction to be executed, thereby reducing the delay for the central processing unit to learn that the operating status of the neural network processor is abnormal.

[0122] 2. The debug and trace functions provided by the monitoring module facilitate debugging and positioning.

[0123] 3. Based on the monitoring algorithm, it is predicted in advance whether the operating status of the neural network processor will be abnormal, so as to make corresponding processing in advance.

[0124] 4. By recording the first target operating status information and / or the second target operating status information, the monitoring algorithm is continuously trained based on the first target operating status information and / or the second target operating status information to continuously improve the accuracy of the prediction results of the monitoring algorithm.

[0125] Corresponding to the aforementioned method for monitoring a neural network processor, the present invention also provides a device for monitoring a neural network processor. Since the device embodiments of the present invention correspond to the aforementioned method embodiments, details not disclosed in the device embodiments can be referred to in the aforementioned method embodiments and will not be further described in this invention.

[0126] Figure 6 A schematic diagram of a device for monitoring a neural network processor according to an embodiment of the present disclosure is shown in FIG. Figure 6 As shown, the device is applied to a neural network processor in a neural network processor monitoring system, including:

[0127] A receiving unit 41 is used to receive a model algorithm task;

[0128] A first acquisition unit 42 is used to acquire first target operating state information during the execution of the model algorithm task;

[0129] The first sending unit 43 is used to send the first target operating state information to the monitoring module, so that the monitoring module sends a first interrupt signal to the central processing unit when determining that the first target operating state information is abnormal. The neural network processor monitoring device provided by the present disclosure adds a detection module to each target computing unit included in the computing cluster module of the neural network processor, and connects each detection module to the monitoring module. When the target computing unit executes the model algorithm task, the first target operating state of the target computing unit is collected based on the detection module, and the first target operating state is sent to the monitoring module. The monitoring module receives the first target operating state. When the monitoring module determines that the first target operating state is abnormal, it triggers a central processing unit interrupt, so that the monitoring module can feedback the operating state of the neural network processor to the central processing unit in real time without waiting for the target instruction to be executed, thereby reducing the delay for the central processing unit to know that the operating state of the neural network processor is abnormal. Figure 7 A schematic diagram of the structure of another apparatus for monitoring a neural network processor provided in an embodiment of the present disclosure, such as Figure 7 As shown, the device also includes:

[0130] A second acquiring unit 44 is configured to acquire second target operating status information, where the second target operating status information is operating status information of a host interface block processor;

[0131] The second sending unit 45 is configured to send the second target operating state information to the control state host interface block processor, so that the control state host interface block processor sends a second interrupt signal to the central processing unit when determining that the second target operating state information is abnormal. Figure 8 This is a schematic diagram of the structure of a device for monitoring a neural network processor provided by an embodiment of the present disclosure. The device is applied to a central processing unit in a neural network processor monitoring system, including:

[0132] A sending unit 51 is configured to send a model algorithm task to a neural network processor, so that the neural network processor obtains first target operating state information during execution of the model algorithm task;

[0133] The execution unit 52 is used to perform an interrupt operation in response to receiving a first interrupt signal sent by the monitoring module, wherein the monitoring module is used to trigger the first interrupt signal when it is determined that the first target operating status information sent by the neural network processor is abnormal.

[0134] Figure 9 This is a schematic diagram of the structure of a device for monitoring a neural network processor provided by an embodiment of the present disclosure, wherein the execution unit 52 is further configured to:

[0135] In response to receiving a second interrupt signal sent by the control state host interface block processor, an interrupt operation is performed, and the control state host interface block processor is used to trigger the second interrupt signal when it is determined that there is an abnormality in the second target operating status information sent by the neural network processor, and the second target operating status information is the operating status information of the host interface block processor.

[0136] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 9 As shown, the device includes:

[0137] a collecting unit 53, configured to collect the first target operating state information received by the monitoring module and / or the second target operating state information received by the control state host interface block processor;

[0138] The saving unit 54 is configured to save the first target running status information and / or the second target running status information to a target server.

[0139] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 9 As shown, the device includes:

[0140] A judging unit 55 is used to judge whether the computing resources of the neural network processor are sufficient;

[0141] a control unit 56 configured to control the host interface block processor to execute a monitoring algorithm when the computing resources of the neural network processor are sufficient, wherein the monitoring algorithm is configured to predict whether an abnormality will occur in the operating state of the neural network processor;

[0142] The control unit 56 is also used to control the central processing unit to run the monitoring algorithm when the computing resources of the neural network processor are insufficient.

[0143] It should be noted that the above explanation of the method embodiment is also applicable to the device of the embodiment of the present disclosure, and the principles are the same, which is no longer limited in the embodiment of the present disclosure.

[0144] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0145] Figure 10A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0146] like Figure 10 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 602 or a computer program loaded from a storage unit 608 into a RAM (Random Access Memory) 603. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An I / O (Input / Output) interface 605 is also connected to the bus 604.

[0147] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0148] Computing unit 601 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 601 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. Computing unit 601 performs the various methods and processes described above, such as the apparatus for monitoring a neural network processor. For example, in some embodiments, the apparatus for monitoring a neural network processor can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by computing unit 601, one or more steps of the method described above can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured as a device for performing the aforementioned neural network processor monitoring in any other appropriate manner (e.g., by means of firmware).

[0149] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0150] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0151] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0152] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0153] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0154] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0155] It's important to note that artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0156] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0157] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A neural network processor monitoring system, characterized in that: include: A computing cluster module includes multiple target computing units, each of which includes a detection module. The computing cluster module is used to receive target instructions, and the target instructions are used to drive the computing cluster module to use the target computing unit to execute the model algorithm task pushed by the central processing unit to the neural network processor; The detection module is configured to obtain first target operating status information of the corresponding target computing unit and send the first target operating status information to the monitoring module; The monitoring module connected to each detection module is used to receive the first target operation status information and trigger the central processing unit to interrupt when determining that the first target operation status information is abnormal.

2. The system according to claim 1, wherein: The central processing unit also includes: A control state host interface block processor connected to the central processing unit and the host interface block processor respectively is used to monitor and obtain the second target operating state information of the host interface block processor, and trigger the central processing unit interrupt when it is determined that the second target operating state information is abnormal, wherein the host interface block processor is a structural component of the neural network processor, and the host interface block processor is used to schedule the model algorithm task, and generate corresponding target instructions according to the model algorithm task, and send the target instructions to the computing cluster module.

3. The system according to claim 1, wherein: The detection module is a monitoring ip block of the target computing unit, and the monitoring ip block is an integrated circuit design component used to monitor and manage the target computing unit; the monitoring module is connected to the central processing unit through an on-chip bus.

4. A method for monitoring a neural network processor, characterized in that: include: Receive model algorithm tasks; During the execution of the model algorithm task, obtaining first target operating state information; The first target operating state information is sent to a monitoring module, so that the monitoring module sends a first interrupt signal to a central processing unit when determining that the first target operating state information is abnormal.

5. The method according to claim 4, characterized in that After receiving the model algorithm task, the method includes: Acquire second target running status information, where the second target running status information is running status information of a host interface block processor; The second target operating state information is sent to the control state host interface block processor, so that the control state host interface block processor sends a second interrupt signal to the central processing unit when determining that the second target operating state information is abnormal.

6. A method for monitoring a neural network processor, characterized in that: include: Sending the model algorithm task to the neural network processor, so that the neural network processor obtains first target operating state information during the execution of the model algorithm task; In response to receiving a first interrupt signal sent by a monitoring module, an interrupt operation is performed, wherein the monitoring module is used to trigger the first interrupt signal when it is determined that the first target operating status information sent by the neural network processor is abnormal.

7. The method according to claim 6, characterized in that After sending the model algorithm task to the neural network processor, the method includes: In response to receiving a second interrupt signal sent by the control state host interface block processor, an interrupt operation is performed, and the control state host interface block processor is used to trigger the second interrupt signal when it is determined that there is an abnormality in the second target operating status information sent by the neural network processor, and the second target operating status information is the operating status information of the host interface block processor.

8. The method according to claim 7, characterized in that After the interrupt operation is performed, the method includes: collecting the first target operating state information received by the monitoring module and / or the second target operating state information received by the control state host interface block processor; The first target running status information and / or the second target running status information are saved to a target server.

9. The method according to claim 8, characterized in that Before sending the model algorithm task to the neural network processor, the method includes: Determining whether the computing resources of the neural network processor are sufficient; If the computing resources of the neural network processor are sufficient, controlling the host interface block processor to run a monitoring algorithm, wherein the monitoring algorithm is used to predict whether an abnormality will occur in the operating state of the neural network processor; If the computing resources of the neural network processor are insufficient, the central processing unit is controlled to run the monitoring algorithm.

10. A device for monitoring a neural network processor, characterized in that: include: A receiving unit, used for receiving model algorithm tasks; A first acquisition unit is used to acquire first target operating state information during the execution of the model algorithm task; The first sending unit is configured to send the first target operating status information to the monitoring module, so that the monitoring module sends a first interrupt signal to the central processing unit when determining that the first target operating status information is abnormal.

11. A device for monitoring a neural network processor, characterized in that: include: a sending unit, configured to send the model algorithm task to the neural network processor, so that the neural network processor obtains first target operating state information during the execution of the model algorithm task; An execution unit is used to perform an interrupt operation in response to receiving a first interrupt signal sent by a monitoring module, wherein the monitoring module is used to trigger the first interrupt signal when determining that the first target operating status information sent by the neural network processor is abnormal.

12. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 4-5 or 6-11.

13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 4-5 or 6-11.

14. A computer program product, characterized in that A computer program is included which, when executed by a processor, implements the method according to any one of claims 4-5 or 6-11.