Software and hardware cooperative real-time performance monitoring device, operation method and computer equipment
By integrating bandwidth and periodic monitoring units within the RVGPNPU IP, the problems of high resource interference and limited functionality in existing technologies are solved. This enables low-overhead, real-time performance monitoring, supports comprehensive indicator detection, and improves the system's real-time performance and detection accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CCORE TECH CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-06-12
Smart Images

Figure CN122195768A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of integrated circuit technology, and in particular to a hardware and software collaborative real-time performance monitoring device, operating method, and computer equipment. Background Technology
[0002] With the widespread adoption of artificial intelligence and edge computing, heterogeneous computing IP cores have become crucial for handling high-performance AI inference tasks. This type of architecture deeply integrates the control capabilities of a RISC-V CPU (Reduced Instruction Set Computer-V Central Processing Unit) with the computing power of a GPNPU (General-Purpose Neural Processing Unit). However, performance optimization and debugging face challenges in practical deployments. System designers need to monitor the internal state of the IP cores (Intellectual Property Cores) in real time, such as the utilization rate of computing units, memory bandwidth bottlenecks, and instruction execution efficiency, to maximize hardware performance and ensure real-time performance.
[0003] Traditional performance monitoring methods typically rely on external tools or software simulations, resulting in significant resource interference and limited functionality, failing to achieve low-overhead, real-time monitoring within the IP core. For example, software monitoring solutions based on external performance analysis tools suffer from microsecond-level latency introduced by the software stack, hindering real-time feedback. Furthermore, frequent interruptions and context switching can interfere with normal task execution, impacting system determinism. While partially integrated solutions based on hardware performance counters reduce software overhead, the counters themselves are limited in functionality, lacking advanced features such as bandwidth detection. Moreover, data aggregation still relies on CPU polling within the instruction set architecture, preventing automatic hardware aggregation and leading to coarse-grained detection and poor real-time performance. Therefore, an effective hardware-software co-operational real-time performance monitoring device is needed to achieve low-overhead real-time monitoring and support comprehensive metric detection. Summary of the Invention
[0004] To address the problems of existing technologies that rely on external tools or software simulation, suffer from significant resource interference and limited functionality, and are unable to achieve low-overhead, real-time monitoring within the IP core, this application mainly provides a hardware and software collaborative real-time performance monitoring device, operating method, and computer equipment based on RVGPNPU.
[0005] To achieve the above objectives, the first technical solution adopted in this application is: a hardware-software collaborative real-time performance monitoring device based on RVGPNPU, comprising: an instruction set architecture central processing unit running firmware; a general-purpose neural network processing unit; and a performance monitoring module integrated within the RVGPNPU, including: a bandwidth monitoring unit that detects the instantaneous bandwidth of the bus data flow in the general-purpose neural network processing unit; an instruction cycle monitoring unit that detects the number of task cycles and the number of multiply-accumulate activation cycles in the general-purpose neural network processing unit; a data aggregation buffer unit that temporarily stores the instantaneous bandwidth, the number of task cycles, and the number of multiply-accumulate activation cycles; and a firmware interface connected to the instruction set architecture central processing unit for firmware access, wherein the firmware obtains the instantaneous bandwidth, the number of task cycles, and the number of multiply-accumulate activation cycles from the data aggregation buffer unit, performs bandwidth bottleneck analysis on the instantaneous bandwidth, and generates a multiply-accumulate utilization rate based on the number of task cycles and the number of multiply-accumulate activation cycles.
[0006] Optionally, the firmware writes bandwidth configuration parameters to the bandwidth monitoring unit via the firmware interface.
[0007] Optionally, a bandwidth parameter configuration subunit is provided, which configures the read byte counter and the write byte counter according to the bandwidth configuration parameters; the read byte counter detects the read bandwidth of the bus data flow; the write byte counter detects the write bandwidth of the bus data flow.
[0008] Optionally, the bandwidth monitoring unit may further include: updating the performance monitoring register, writing the read bandwidth and write bandwidth, and setting the interrupt flag bit according to the instantaneous bandwidth and the bandwidth threshold in response to determining that a configured bandwidth threshold has been determined.
[0009] Optionally, the hardware-software co-processor real-time performance monitoring device also includes a hybrid scheduling unit connected between the general-purpose neural network processing unit and the instruction set architecture central processing unit.
[0010] Optionally, the firmware writes cycle configuration parameters to the instruction cycle monitoring unit via memory mapping.
[0011] Optionally, the instruction cycle monitoring unit includes: a cycle parameter configuration subunit, which configures the parameters of the task cycle monitor and the multiply-accumulate unit monitor according to the cycle configuration parameters; a task cycle counter, which controls the start and stop of the hybrid scheduling unit to determine the number of task cycles consumed by a single task; and a multiply-accumulate unit activity counter, which summarizes the multiply-accumulate unit activation signals of each computing processing unit in the computing processing unit array to determine the number of multiply-accumulate activation cycles during task execution.
[0012] The second technical solution adopted in this application is: an operation method for a hardware-software collaborative real-time performance monitoring device based on RVGPNPU, comprising: a bandwidth detection step, wherein the instantaneous bandwidth of the bus data flow in the general neural network processing unit is detected by the bandwidth monitoring unit included in the performance monitoring module; an instruction cycle monitoring step, wherein the task cycle number and multiply-accumulate activation cycle number of the general neural network processing unit are detected by the instruction cycle monitoring unit included in the performance monitoring module; a temporary storage step, wherein the instantaneous bandwidth, task cycle number, and multiply-accumulate activation cycle number are temporarily stored by the data aggregation buffer unit included in the performance monitoring module; and an analysis and generation step, wherein the firmware obtains the instantaneous bandwidth, task cycle number, and multiply-accumulate activation cycle number from the data aggregation buffer unit, performs bandwidth bottleneck analysis on the instantaneous bandwidth, and generates a multiply-accumulate utilization rate based on the task cycle number and multiply-accumulate activation cycle number, wherein the firmware interface is connected to the instruction set architecture central processing unit for firmware access.
[0013] Optionally, the bandwidth detection steps include: a bandwidth parameter configuration step, which configures the read byte counter and the write byte counter according to the bandwidth configuration parameters written in the firmware; a read bandwidth detection step, which detects the read bandwidth of the bus data traffic; and a write bandwidth detection step, which detects the read bandwidth of the bus data traffic.
[0014] The third technical solution adopted in this application is: a computer device, including a memory, an instruction set architecture central processing unit, and a computer program stored in the memory, wherein the instruction set architecture central processing unit executes the computer program to implement the operation method of the hardware and software collaborative real-time performance monitoring device in the second solution.
[0015] The beneficial effects that the technical solution of this application can achieve are as follows: When applied, the technical solution of this application integrates performance detection function within the IP of RVGPNPU, with software as the main component and hardware as the auxiliary component, eliminating the dependence on external tools, realizing nanosecond-level latency monitoring, achieving low-overhead real-time monitoring; and expanding the detection range, including bandwidth and period, supporting comprehensive indicator detection. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of a specific embodiment of a hardware and software collaborative real-time performance monitoring device according to this application; Figure 2This is an overall system diagram of a specific embodiment of a hardware and software collaborative real-time performance monitoring device according to this application; Figure 3 This is a flowchart of a specific embodiment of the operation method of a hardware and software collaborative real-time performance monitoring device according to this application; Figure 4 This is a flowchart of a specific embodiment of the bandwidth detection step in the operation method of a hardware and software collaborative real-time performance monitoring device of this application.
[0018] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0019] The preferred embodiments of this application will now be described in detail with reference to the accompanying drawings, so that the advantages and features of this application can be more easily understood by those skilled in the art, thereby providing a clearer and more definite definition of the scope of protection of this application.
[0020] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0021] The technical solution of this application and how it solves the above-mentioned technical problems will be described in detail below with specific embodiments. The specific embodiments described below can be combined with each other to form new embodiments. The same or similar ideas or processes described in one embodiment may not be repeated in other embodiments.
[0022] Figure 1 This is a schematic diagram of a specific embodiment of a hardware and software collaborative real-time performance monitoring device according to this application.
[0023] Figure 2 This is an overall system diagram of a specific embodiment of a hardware and software collaborative real-time performance monitoring device according to this application.
[0024] Figure 1 The hardware and software collaborative real-time performance monitoring device shown includes: an instruction set architecture central processing unit 101, which runs firmware.
[0025] In one specific embodiment of this application, the instruction set architecture CPU 101 is a RISC-V Core. The instruction set architecture CPU 101 is responsible for executing the operating system, network protocol stack, and top-level business logic. The instruction set architecture CPU 101 includes a general-purpose register file, a CSR register (Control and Status Register), a PMCR register (Performance Monitor Cycle Register), a PMBR register (Performance Monitor Bandwidth Register), and a custom instruction extension interface. The firmware can be software running on the instruction set architecture CPU 101. The firmware can configure hardware, periodically read data, and perform summary analysis through a memory-mapped interface. The firmware implements intelligent strategies, such as dynamically adjusting the detection frequency.
[0026] Figure 1 The hardware and software collaborative real-time performance monitoring device shown includes: a general neural network processing unit 102, which performs tensor operations.
[0027] In one specific embodiment of this application, the general neural network processing unit 102 is a GPNPU array, which includes multiple processing elements (PEs), register file groups, and internal DMA (Direct Memory Access), and is responsible for performing high-throughput tensor operations.
[0028] Figure 1 The hardware-software collaborative real-time performance monitoring device shown includes: a performance monitoring module, which is integrated inside the RVGPNPU, and includes: The bandwidth monitoring unit 103 detects the instantaneous bandwidth of the bus data flow in the general neural network processing unit.
[0029] In one specific embodiment of this application, the bandwidth monitoring unit 103 is a hardware module that monitors the bus data traffic in the general neural network processing unit and calculates the instantaneous bandwidth. The bandwidth monitoring accuracy is one sample every 100 ns.
[0030] Optionally, the firmware writes bandwidth configuration parameters to the bandwidth monitoring unit via the firmware interface. Specifically, the firmware can write configuration parameters to the bandwidth monitoring unit through the bandwidth performance monitoring register (PMBR) of the firmware interface. The configuration parameters mainly include the monitoring enable flag, read / write channel monitoring enable, and time window length (window_size, unit: AXI (Advanced eXtensible Interface) clock cycle). The AXI clock cycle refers to one complete pulse cycle of the clock signal in the AXI bus protocol. It is the smallest time unit for AXI bus data transmission and timing synchronization, determining the bus's transmission rate and data throughput.
[0031] Optionally, the bandwidth monitoring unit 103 includes a bandwidth parameter configuration subunit, which configures the read byte counter and write byte counter according to the bandwidth configuration parameters. Specifically, the bandwidth monitoring unit 103 can receive configuration parameters, reset internal counters and state machines, start an AXI clock-based window counter, and clear the read byte counter and write byte counter.
[0032] The read byte counter detects the read bandwidth of the bus data flow. Specifically, first, it continuously monitors the AXI read data channel signals: when RVALID (read data valid signal) == 1 and RREADY (read data ready signal) == 1, it indicates that a valid read beat has been completed. By default, the BandWidth (assumed to be 128-bit for simplicity) is fully valid (16 bytes), and this is accumulated in read_byte_cnt. If read monitoring is enabled, the accumulation is performed; otherwise, it is skipped. Then, it checks whether the window counter meets a preset counting condition. If it determines that the window counter does not meet the preset counting condition, it continues to monitor the AXI read data channel signals. If it determines that the window counter meets the preset counting condition, it stops monitoring the AXI read data channel signals. The preset counting condition can be whether the window counter value has reached a preset value. Finally, the ratio of the accumulated read_byte_cnt to the product of the preset value and the clock cycle can be determined as the read bandwidth.
[0033] The write byte counter detects the write bandwidth of the bus data flow. Specifically, first, it continuously monitors the AXI write data channel signals: when WVALID (write data valid signal) == 1 and WREADY (write data ready signal) == 1, it indicates that a valid write beat (the smallest data transmission unit of the AXI bus) has been completed. The number of valid bytes in the current beat is calculated based on the WSTRB (write strobe signal). If write monitoring is enabled, the number of valid bytes is accumulated to write_byte_cnt. Then, it checks whether the window counter meets a preset counting condition. In response to determining that the window counter does not meet the preset counting condition, it continues to monitor the AXI write data channel signals. In response to determining that the window counter meets the preset counting condition, it stops monitoring the AXI write data channel signals. The preset counting condition can be whether the value of the window counter has reached a preset value. Finally, the ratio of the accumulated write_byte_cnt to the product of the preset value and the clock cycle can be determined as the write bandwidth.
[0034] Optionally, the bandwidth monitoring unit 103 further includes: updating the performance monitoring register, which writes the read bandwidth and write bandwidth, and setting an interrupt flag bit based on the instantaneous bandwidth and the bandwidth threshold in response to determining that a bandwidth threshold is configured. Specifically, updating the performance monitoring register writes the generated write bandwidth or read bandwidth to a dedicated read-only register. For example, the calculation results are written to the dedicated read-only register: Read_BW (Read Bandwidth) is written to READ_BW_REG (read bandwidth register); Write_BW (Write Bandwidth) is written to WRITE_BW_REG (write bandwidth register); Read_BW + Write_BW is written to TOTAL_BW_REG (total bandwidth register). Furthermore, if a bandwidth threshold is configured, the current instantaneous bandwidth can be compared with the bandwidth threshold, and the current + Write_BW can be written to TOTAL_BW_REG (total bandwidth register). Also, if a bandwidth threshold is configured, an interrupt flag bit (such as BW_OVERRUN_FLAG) can be set when the current instantaneous bandwidth is less than or equal to the bandwidth threshold.
[0035] The bandwidth monitoring unit 103 may further include a reset subunit, which clears the read byte counter and write byte counter and resets the window counter. Specifically, it clears the read byte count value and / or the write byte count value. The window counter is reset to 0.
[0036] The instruction cycle monitoring unit 104 detects the number of task cycles and the number of multiply-accumulate activation cycles of the general neural network processing unit.
[0037] In one specific embodiment of this application, the instruction cycle monitoring unit 104 can be an integrated performance counter that detects the number of task cycles and multiply-accumulate activation cycles of the general neural network processing unit. The counter precision for instruction cycle monitoring is in units of one clock cycle. Thus, the newly added hardware units (such as bandwidth detectors and performance counters) are responsible for low-level data acquisition, operating in the RISC-V clock domain, avoiding additional latency. The instruction cycle monitoring unit covers statistical events related to GPNPU operator execution and MAC (Multiply-Accumulate) usage, while the bandwidth detection unit directly monitors the data throughput rate events of the physical bus. It supports multi-dimensional metrics such as bandwidth and cycle time, and the hardware counters provide instruction-level precision, overcoming the limitation of single-function solutions in existing approaches.
[0038] Optionally, the hardware-software collaborative real-time performance monitoring device also includes a hybrid scheduling unit, which connects the general-purpose neural network processing unit (GNRF) and the instruction set architecture (IPA) CPU. Specifically, the hybrid scheduling unit is located between the RISC-V core and the GPNPU array, and is used to realize data transfer between the GNRF and the IPA CPU, handling instruction parsing, dependency analysis, task distribution, and state synchronization. Thus, the performance monitoring module can utilize the task state information from the hybrid scheduling unit to achieve context-aware monitoring in conjunction with task scheduling. Furthermore, based on the existing RVGPNPU architecture, the hybrid scheduling unit and memory interface are reused, reducing implementation costs.
[0039] Optionally, the firmware writes cycle configuration parameters to the instruction cycle monitoring unit via memory mapping. Specifically, the firmware sets `enable_cycle_monitor` (clock cycle monitoring enable variable) to 1 to enable task cycle monitoring; and sets `enable_mac_monitor` (multiply-accumulate activation cycle monitoring enable variable) to 1 to enable multiply-accumulate activation cycle monitoring. The cycle configuration parameters can include enabling task cycle monitoring and enabling multiply-accumulate activation cycle monitoring.
[0040] Optionally, the instruction cycle monitoring unit includes a cycle parameter configuration subunit, which configures the task cycle monitor and the multiply-accumulate unit monitor according to the cycle configuration parameters. Specifically, when the hybrid scheduling unit sends a task to the GPNPU, it simultaneously sends a task_start signal to the instruction cycle monitoring unit, along with the task ID. The cycle parameter configuration subunit can write the task ID to TASK_ID_REG and clear the task cycle counter and the multiply-accumulate unit activity counter.
[0041] The task cycle counter, controlled by the hybrid scheduling unit, determines the number of task cycles consumed by a single task. Specifically, the task cycle counter increments every GPNPU clock cycle.
[0042] The multiply-accumulate unit activity counter summarizes the multiply-accumulate unit activation signals of each computing processing unit in the computing processing unit array to determine the number of multiply-accumulate activation cycles during task execution. Specifically, the multiply-accumulate unit activity counter accumulates each clock cycle based on the mac_active_pulse signal (MAC unit active pulse signal).
[0043] After completing the task, the GPNPU sends a stop (task_done) signal to the hybrid scheduling unit. The hybrid scheduling unit forwards this signal to the instruction cycle monitoring unit, triggering: latching the task cycle counter value to the task cycle counter; latching the multiply-accumulate unit activity counter value to the multiply-accumulate unit activity counter; and setting the data_ready_flag (data ready flag) to 1.
[0044] Optionally, the instruction cycle monitoring unit may also include a task identifier register, which receives a unique task identifier injected by the hybrid scheduling unit.
[0045] Optionally, the instruction cycle monitoring unit may also include: a cycle performance monitoring register, which enables task cycle monitoring, enables multiply-accumulate unit monitoring, and registers a data ready flag and a current task identifier; and a dedicated read-only register, which registers the number of task execution cycles and the number of multiply-accumulate unit activation cycles.
[0046] The data aggregation buffer unit 105 temporarily stores instantaneous bandwidth, task cycle count, and multiply-accumulate activation cycle count.
[0047] In one specific embodiment of this application, the data aggregation buffer unit 105 is a shared memory region used to temporarily store the raw monitoring data monitored by the bandwidth monitoring unit 103 and the instruction cycle monitoring unit 104, namely, instantaneous bandwidth, task cycle count, and multiply-accumulate activation cycle count. A fixed area is allocated in the shared memory, and the hardware unit directly writes data via DMA. The firmware obtains the instantaneous bandwidth, task cycle count, and multiply-accumulate activation cycle count temporarily stored in the data aggregation buffer unit by directly manipulating the physical address. Thus, monitoring data can be transferred through the internal shared memory of the IP, and the firmware can directly access it without bus arbitration or copying. This avoids the operating system protocol stack, reduces intermediate steps, and lowers the detection latency from milliseconds to microseconds, which is superior to software-dependent solutions, achieving low-latency real-time monitoring.
[0048] Firmware interface 106 is connected to the instruction set architecture central processing unit for firmware access. The firmware obtains instantaneous bandwidth, task cycle count, and multiply-accumulate activation cycle count from the data aggregation buffer unit, performs bandwidth bottleneck analysis on the instantaneous bandwidth, and generates multiply-accumulate utilization based on the task cycle count and multiply-accumulate activation cycle count.
[0049] In one specific embodiment of this application, the firmware interface 106 is connected to the instruction set architecture central processing unit and exposes its configuration and status through the PMCR and PMBR registers for firmware access. The PMCR and PMBR are used to enable firmware configuration monitoring data, read status flags, enable MAC utilization monitoring, instruction execution cycle detection, and bandwidth monitoring bit flags.
[0050] The firmware periodically checks the temporary storage status of the data aggregation buffer module. In response to the detection of temporary storage of instantaneous bandwidth, task cycles, and multiply-accumulate activation cycles, it retrieves these data from the data aggregation buffer unit. It then performs bandwidth bottleneck analysis on the instantaneous bandwidth and generates a multiply-accumulate utilization rate based on the task cycles and multiply-accumulate activation cycles. Specifically, during the initialization phase, the firmware configures the performance monitoring module through the PMCR and PMBR registers, setting the monitoring objects for bandwidth detection. These primarily include the real-time data throughput rate of the AXI bus and memory ports, and the event types for periodic monitoring (mainly two indicators: GPNPU instruction execution cycle and GPNPU instruction idle cycle) and MAC utilization. At this time, the bandwidth monitoring unit and the cycle monitoring unit load their configurations and enter a standby state. Then, the firmware enables the bandwidth monitoring unit and the instruction cycle monitoring unit, starting background monitoring. The bandwidth monitoring unit and the instruction cycle monitoring unit begin automatically collecting data, but the firmware does not process it immediately. The firmware sends a monitoring enable signal through the PMCR and PMBR registers to start the bandwidth monitoring unit and the instruction cycle monitoring unit. The bandwidth monitoring unit and instruction cycle monitoring unit begin background monitoring: the bandwidth monitoring unit monitors data traffic, and the instruction cycle monitoring unit monitors the number of task cycles and the number of multiply-accumulate activation cycles. The firmware enters a low-power wait state, awaiting an interrupt trigger. During the runtime monitoring phase, firstly, the bandwidth monitoring unit and instruction cycle monitoring unit continuously monitor bandwidth and cycle data. The bandwidth monitoring unit statistically analyzes the data transmission volume within a specified time window, calculates the instantaneous bandwidth, and calculates the number of task cycles and the number of multiply-accumulate activation cycles. Then, when data is ready or a threshold is reached, the bandwidth monitoring unit and instruction cycle monitoring unit automatically write the data packet (including timestamp and event identifier) to the data aggregation buffer via DMA. The writing process does not require intervention from the instruction set architecture CPU. Afterward, the firmware periodically (e.g., every millisecond) checks the buffer status via lightweight interrupts or polling the PMCR and PMBR flags. Finally, it detects the arrival of new data. The firmware reads the performance monitoring register data and performs preliminary aggregation based on the timestamp and corresponding event type (e.g., calculating the average bandwidth, MAC utilization, and number of instruction cycles during operator execution). The summarized results are temporarily stored in the local cache. In the data analysis and transmission phase, the firmware first performs calculations and analysis on the aggregated data according to application requirements. Then, the resulting data is transmitted to an external system or log file via storage. The firmware uses standard memory sharing technology to copy data to a designated area or sends it to a remote monitoring terminal via a network protocol stack. Finally, the firmware can dynamically adjust the RVGPNPU operating parameters based on the analysis results (e.g., limiting task priorities through a hybrid scheduling unit) to achieve closed-loop optimization. Thus, in the entire process, hardware handles high-frequency data acquisition, while software handles low-frequency intelligent processing, balancing real-time performance and flexibility. Monitoring data flow and task execution pipeline run in parallel to ensure minimal interference.Furthermore, the data aggregation buffer enables zero-copy transfer, reducing bus contention; the firmware adopts a lightweight triggering mechanism, decoupling the monitoring process from task execution, with the hardware running automatically and the firmware only intermittently intervening to ensure that critical tasks are not affected.
[0051] Optionally, the firmware includes a task identifier-operator type mapping table. In response to detecting a new task submitted by the hybrid scheduling unit, the firmware writes the task identifier corresponding to the new task into the scheduling instruction. Specifically, the firmware maintains a task ID-operator type mapping table (e.g., ID=0x1 corresponds to a convolutional layer, ID=0x2 corresponds to a fully connected layer, etc.). When a new task is submitted to the hybrid scheduling unit, the firmware simultaneously writes the task ID into the scheduling instruction.
[0052] Optionally, the firmware, after detecting data readiness via polling or lightweight interrupts, sequentially reads data from the task identifier register, task cycle monitor, multiply-accumulate unit monitor, write channel transaction monitoring subunit, and read channel transaction monitoring subunit. Based on the data read from the task cycle monitor and multiply-accumulate unit monitor, it generates the utilization rate of the multiply-accumulate unit within a specified time window. Furthermore, based on the data read from the write channel transaction monitoring subunit and read channel transaction monitoring subunit, it performs bandwidth bottleneck analysis to determine if the current bandwidth is a performance bottleneck. Specifically, the instruction set architecture CPU writes to the PMCR register via memory mapping: setting enable_cycle_monitor (clock cycle monitoring enable variable) = 1 to enable task cycle monitoring; setting enable_mac_monitor (multiply-accumulate unit monitoring enable variable) = 1 to enable MAC utilization monitoring; the performance monitoring module enters standby mode upon receiving the configuration. During task startup synchronization, when the hybrid scheduling unit sends a task to the GPNPU, it simultaneously sends a task_start signal to the Cycle / MAC (task cycle / multiply-accumulate) monitoring unit, along with the task ID. The monitoring unit immediately: writes the task ID to TASK_ID_REG (task identifier register); clears the task cycle counter and MAC activity counter; and starts accumulating both counters. During parallel data acquisition, in the Cycle / MAC path: the task cycle counter increments each GPNPU clock cycle; the MAC activity counter accumulates each clock cycle based on the mac_active_pulse signal (MAC unit active pulse signal). During event completion and data latching, in the Cycle / MAC path: after the GPNPU completes the task, it sends a task_done signal (task completion status signal) to the hybrid scheduling unit. The scheduling unit forwards this signal to the monitoring unit, triggering: latching the task cycle counter value to TASK_CYCLE_REG (task cycle register); latching the MAC activity counter value to MAC_ACTIVE_CYCLES_REG (MAC active cycle counter register); and setting data_ready_flag (data ready flag) to 1. When the firmware reads register data, after detecting that data_ready_flag=1 through polling or a lightweight interrupt, it reads the following in sequence: TASK_ID_REG, TASK_CYCLE_REG, MAC_ACTIVE_CYCLES_REG; READ_BW_REG, WRITE_BW_REG. When calculating core performance metrics, the firmware performs the following calculation: MAC Utilization: MAC Utilization = (TASK_CYCLE_REG / MAC_ACTIVE_CYCLES_REG) × 100%.Bandwidth bottleneck analysis: By combining task IDs (such as convolutional layers which typically have high bandwidth requirements), it can be determined whether the current bandwidth is becoming a performance bottleneck. Therefore, the software firmware can dynamically configure detection strategies to adapt to different scenarios (such as debug mode or production monitoring), while the hardware ensures basic efficiency.
[0053] As an example, this application is applied to the performance optimization of edge AI devices, where the memory bandwidth usage of the GPNPU is monitored in real time within the edge AI device. Specifically, firstly, the bandwidth monitoring unit is configured through the PMCR register, setting the monitoring objects to the AXI bus and memory ports. Then, the bandwidth monitoring unit continuously monitors data traffic and calculates the instantaneous bandwidth. Subsequently, when the bandwidth exceeds a threshold (e.g., 80%), the bandwidth monitoring unit triggers an interrupt. Finally, the RISC-V CPU reads the monitoring data and dynamically adjusts the task scheduling strategy, prioritizing tasks with low bandwidth requirements.
[0054] As another example, this application is applied to latency optimization in autonomous driving perception systems. In these systems, high-latency instructions are identified through cycle monitoring. Specifically, first, an instruction cycle monitoring unit is configured via the PMCR register, setting the monitoring event to the GPNPU instruction execution cycle. Then, the instruction cycle monitoring unit calculates the number of execution cycles corresponding to the instruction. Next, the RISC-V CPU reads the monitoring data, identifies the type of high-latency operator, and the specific quantized execution data. Finally, the code path corresponding to the high-latency instruction is optimized, or hardware accelerator optimization is triggered.
[0055] The beneficial effects that the technical solution of this application can achieve are as follows: When applied, the technical solution of this application integrates performance detection function inside the RVGPNPU IP, with software as the main component and hardware as the auxiliary component, eliminating the dependence on external tools, realizing nanosecond-level latency monitoring, achieving low-overhead real-time monitoring; and expanding the detection range, including bandwidth and period, supporting comprehensive indicator detection.
[0056] Figure 3 This paper illustrates a specific implementation of a hardware-software collaborative real-time performance monitoring device for secure data.
[0057] exist Figure 3In the specific implementation shown, the operation method of the hardware-software collaborative real-time performance monitoring device mainly includes: a bandwidth detection step S301, in which the instantaneous bandwidth of the bus data flow in the general neural network processing unit is detected by the bandwidth monitoring unit included in the performance monitoring module; an instruction cycle monitoring step S302, in which the task cycle number and multiply-accumulate activation cycle number of the general neural network processing unit are detected by the instruction cycle monitoring unit included in the performance monitoring module; a temporary storage step S303, in which the instantaneous bandwidth, task cycle number, and multiply-accumulate activation cycle number are temporarily stored by the data aggregation buffer unit included in the performance monitoring module; and an analysis and generation step S304, in which the firmware obtains the instantaneous bandwidth, task cycle number, and multiply-accumulate activation cycle number from the data aggregation buffer unit, performs bandwidth bottleneck analysis on the instantaneous bandwidth, and generates a multiply-accumulate utilization rate based on the task cycle number and multiply-accumulate activation cycle number. The firmware interface is connected to the instruction set architecture central processing unit for firmware access.
[0058] In one specific embodiment of this application, such as Figure 4 As shown, the bandwidth detection step S301 includes: a bandwidth parameter configuration step, which configures the read byte counter and write byte counter according to the bandwidth configuration parameters written in the firmware; a read bandwidth detection step, which detects the read bandwidth of the bus data traffic; and a write bandwidth detection step, which detects the read bandwidth of the bus data traffic. Specifically, the bandwidth detection step S301 includes: a bandwidth parameter configuration step S3011, a time window counter startup step S3012, a write channel transaction monitoring step S3013, a read channel transaction monitoring step S3014, a time window counting step S3015, and an instantaneous bandwidth generation step S3016.
[0059] The operation method of the hardware-software collaborative real-time performance monitoring device provided in this application can be used to execute the hardware-software collaborative real-time performance monitoring device described in any of the above embodiments. The implementation principle and technical effect are similar, and will not be repeated here.
[0060] In one specific embodiment of this application, a computer device includes a memory, an instruction set architecture central processing unit (CPU), and a computer program stored in the memory. The CPU executes the computer program to implement the operation method of the hardware-software collaborative real-time performance monitoring device described in the above embodiments.
[0061] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0062] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0063] The above are merely embodiments of this application and do not limit the scope of this patent application. Any equivalent structural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of this application.
Claims
1. A hardware and software collaborative real-time performance monitoring device based on RVGPNPU, characterized in that, include: Instruction set architecture central processing unit, which runs firmware; General-purpose neural network processing unit; The performance monitoring module, integrated within the RVGPNPU, includes: A bandwidth monitoring unit that detects the instantaneous bandwidth of the bus data traffic in the general neural network processing unit; The instruction cycle monitoring unit detects the number of task cycles and the number of multiply-accumulate activation cycles of the general neural network processing unit. A data aggregation buffer unit temporarily stores the instantaneous bandwidth, the number of task cycles, and the number of multiply-accumulate activation cycles; A firmware interface is connected to the instruction set architecture central processing unit and accessed by the firmware. The firmware obtains the instantaneous bandwidth, the number of task cycles, and the number of multiply-accumulate activation cycles from the data aggregation buffer unit, performs bandwidth bottleneck analysis on the instantaneous bandwidth, and generates a multiply-accumulate utilization rate based on the number of task cycles and the number of multiply-accumulate activation cycles.
2. The hardware and software collaborative real-time performance monitoring device according to claim 1, characterized in that, The firmware writes bandwidth configuration parameters to the bandwidth monitoring unit through the firmware interface.
3. The hardware and software collaborative real-time performance monitoring device according to claim 2, characterized in that, The bandwidth monitoring unit includes: The bandwidth parameter configuration subunit configures the read byte counter and the write byte counter according to the bandwidth configuration parameters. The read byte counter detects the read bandwidth of the bus data flow; The write byte counter detects the write bandwidth of the bus data flow.
4. The hardware and software collaborative real-time performance monitoring device according to claim 3, characterized in that, The bandwidth monitoring unit also includes: Update the performance monitoring register, writing the read bandwidth and the write bandwidth, and in response to determining that a bandwidth threshold is configured, set the interrupt flag bit based on the instantaneous bandwidth and the bandwidth threshold.
5. The hardware and software collaborative real-time performance monitoring device according to claim 1, characterized in that, The hardware-software collaborative real-time performance monitoring device also includes a hybrid scheduling unit, which is connected between the general neural network processing unit and the instruction set architecture central processing unit.
6. The hardware and software collaborative real-time performance monitoring device according to claim 5, characterized in that, The firmware writes cycle configuration parameters to the instruction cycle monitoring unit via memory mapping.
7. The hardware and software collaborative real-time performance monitoring device according to claim 6, characterized in that, The instruction cycle monitoring unit includes: The cycle parameter configuration subunit configures the parameters of the task cycle monitor and the multiply-accumulate unit monitor according to the cycle configuration parameters. The task cycle counter is started and stopped by the hybrid scheduling unit to determine the number of task cycles consumed by a single task. The multiply-accumulate unit activity counter summarizes the multiply-accumulate unit activation signals of each computing processing unit in the computing processing unit array to determine the number of multiply-accumulate activation cycles during task execution.
8. An operation method for a hardware and software collaborative real-time performance monitoring device based on RVGPNPU, characterized in that, include: The bandwidth detection step involves detecting the instantaneous bandwidth of the bus data traffic in the general neural network processing unit through the bandwidth monitoring unit included in the performance monitoring module. The instruction cycle monitoring step involves detecting the number of task cycles and the number of multiply-accumulate activation cycles of the general neural network processing unit through the instruction cycle monitoring unit included in the performance monitoring module. The temporary storage step involves temporarily storing the instantaneous bandwidth, the number of task cycles, and the number of multiply-accumulate activation cycles through the data aggregation buffer unit included in the performance monitoring module. In the analysis and generation steps, the firmware obtains the instantaneous bandwidth, the number of task cycles, and the number of multiply-accumulate activation cycles from the data aggregation buffer unit, performs bandwidth bottleneck analysis on the instantaneous bandwidth, and generates a multiply-accumulate utilization rate based on the number of task cycles and the number of multiply-accumulate activation cycles. The firmware interface is connected to the instruction set architecture central processing unit for the firmware to access.
9. The operation method of the hardware and software collaborative real-time performance monitoring device according to claim 8, characterized in that, The bandwidth detection step includes: The bandwidth parameter configuration step involves configuring the read byte counter and write byte counter according to the bandwidth configuration parameters written in the firmware. The read bandwidth detection step detects the read bandwidth of the bus data traffic; The write bandwidth detection step detects the read bandwidth of the bus data traffic.
10. A computer device comprising a memory, an instruction set architecture central processing unit, and a computer program stored in the memory, characterized in that, The instruction set architecture central processing unit executes the computer program to implement the operation method of the hardware and software collaborative real-time performance monitoring device as described in any one of claims 8-9.