A reconfigurable processor system and method of instruction performance analysis
Patent Information
- Application Number
- CN202610770616.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]现有方案通常依赖估算、软件模拟或粗略统计一段时间内工作单元的工作状态,存在无法在可变时钟域下保证时间统计的准确性、难以精确获取每条指令的实际执行时间以及难以低成本的同时支持在顺序、乱序执行环境下准确追踪指令
[0025]本发明实施例提供的一种可重构处理器系统及指令性能分析方法,所有时间统计均在独立且频率固定的专用时钟域中完成,从而隔离了处理器核心可变时钟域的影响;控制模块为分发的每一条指令分配全局唯一的性能分析ID,能够在全系统中对其进行无歧义地追踪与关联,能够在不显著增加系统面积和功耗的前提下,实现对可重构处理器内各工作单元指令执行时间的精确统计、统一管理和高效上报。
Smart Images

Figure CN122817154A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of processor technology, and more specifically to instruction performance analysis methods, systems, chips, boards, and electronic devices for reconfigurable processors. Background Technology
[0002] Analyzing the instruction performance of reconfigurable processors can help chip designers, compiler developers, and others identify performance bottlenecks in the working units of reconfigurable processors, enabling targeted optimization. With the widespread application of reconfigurable processors in artificial intelligence inference and training, their internal structure and functions are becoming increasingly complex.
[0003] Reconfigurable processors typically include a centralized control module for instruction fetching, decoding, and dispatching; and multiple parallel working units, such as tensor computation, vector computation, scalar computation, cache management, loading, and storing, each executing different types of instructions. Because the execution times of each unit differ significantly, the execution time of each unit needs to be independently calculated; multiple instructions can also be executed in parallel within the same working unit, and the completion order may be out of order. Furthermore, the processor may operate in a dynamically variable clock domain; after completing the time statistics for each unit, the system also needs to perform unified aggregation and comprehensive analysis of the instruction execution time data for all working units.
[0004] Existing solutions typically rely on estimation, software simulation, or rough statistics of the working state of work units over a period of time. These solutions cannot guarantee the accuracy of time statistics under variable clock domains, are difficult to accurately obtain the actual execution time of each instruction, and are difficult to support accurate instruction tracking in sequential and out-of-order execution environments at low cost. Summary of the Invention
[0005] To address at least one of the technical problems raised in the background section, this application provides a reconfigurable processor system and instruction performance analysis method, which can achieve accurate statistics, unified management, and efficient reporting of instruction execution time of each working unit within the reconfigurable processor without significantly increasing system area and power consumption.
[0006] In a first aspect, embodiments of this application provide a reconfigurable processor system, including a reconfigurable processor, a data export module, and a performance analysis module. The reconfigurable processor includes a control module, multiple parallel working units, and an on-chip storage module, wherein: The control module is configured to generate a performance analysis enable signal and assign a globally unique performance analysis ID to each instruction before distributing the instructions to the work unit. The performance analysis module includes performance analysis units corresponding to each of the working units, wherein each of the performance analysis units is configured as follows: Based on the performance analysis enable signal sent by the control module, it is determined whether the performance analysis function of the working unit is enabled; in response to determining that the performance analysis function of the working unit is enabled, based on the instruction information of each instruction sent by the working unit and the instruction execution mode of the working unit, the corresponding instruction performance analysis operation is performed on the working unit in a fixed clock domain to obtain the performance analysis result of the working unit, wherein the instruction information includes the performance analysis ID; The data export module is configured to: read the performance analysis results from the cache queue corresponding to each working unit based on a polling strategy, and export the performance analysis results to the on-chip storage module through the on-chip bus interface.
[0007] In some optional embodiments of this example, the working unit includes m working sub-units, and the performance analysis unit includes n performance analysis sub-units, where n equals m and n is a natural number greater than 0. Each performance analysis sub-unit is configured as follows: In response to the instruction execution mode being a sequential execution mode, based on the instruction information sent by the working unit, a sequential instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information further includes an instruction start signal and an instruction end signal; In response to the instruction execution mode being out-of-order execution mode, based on the instruction information of each instruction sent by the working unit, an out-of-order instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information also includes the instruction's initial ID, instruction start signal, and instruction end signal.
[0008] In some optional embodiments of this example, when n is greater than 1, the performance analysis unit is further configured as follows: The effective output data of each performance analysis subunit is collected sequentially using a polling strategy, and the effective output data is sent to the cache queue corresponding to the working unit.
[0009] In some optional embodiments of this example, in response to the instruction execution mode being a sequential execution mode, the performance analysis subunit is further configured as follows: In response to receiving the instruction start signal, the current address to be written is determined based on the previous write address, and the instruction start time is determined based on a pre-set first counter; Based on the current address to be written, the performance analysis ID, and the instruction start time, update the first execution time statistics table corresponding to the working subunit; In response to receiving the instruction end signal, the current address to be read is determined based on the previous read address; Based on the current address to be read, the first execution time statistics table is read, and the read valid output data is sent to the cache queue corresponding to the work unit.
[0010] In some optional embodiments of this example, the first execution time statistics table includes an ID field, a start time field, an execution time field, and an instruction status field, and the performance analysis subunit is further configured as follows: Based on the current address to be written, write the performance analysis ID into the number field of the first execution time statistics table, write the instruction start time into the start time field, and write 1 into the instruction status field; In response to the instruction status field being set to 1, a second counter is started to keep time until the instruction end signal is received, the instruction execution time is obtained, and the instruction execution time is written into the execution time field.
[0011] In some optional embodiments of this example, the minimum depth of the first execution time statistics table is set to the maximum number of instruction pipelines; the maximum depth of the first execution time statistics table is set to the maximum number of instructions that the control module can distribute.
[0012] In some optional embodiments of this example, in response to the instruction execution order being out-of-order execution mode, the performance analysis subunit is further configured as follows: In response to receiving the instruction start signal, the current address to be written is determined based on the initial ID of the instruction corresponding to the instruction start signal, and the instruction start time is determined based on a pre-set first counter; Based on the current address to be written, the performance analysis ID, and the instruction start time, update the second execution time statistics table corresponding to the working subunit; In response to receiving the instruction end signal, the current address to be read is determined based on the initial ID of the instruction corresponding to the instruction end signal; Based on the current address to be read, the second execution time statistics table is read, and the read valid output data is sent to the cache queue corresponding to the work unit.
[0013] In some optional embodiments of this example, the second execution time statistics table includes an ID field, a start time field, an execution time field, and an instruction status field. Updating the second execution time statistics table corresponding to the work unit based on the current address to be written, the performance analysis ID, and the instruction start time includes: Based on the current address to be written, write the performance analysis ID into the number field of the second execution time statistics table, write the instruction start time into the start time field, and write 1 into the instruction status field; In response to the instruction status field being set to 1, a second counter is started to keep time until the instruction end signal is received, the instruction execution time is obtained, and the instruction execution time is written into the execution time field.
[0014] In some alternative embodiments of this example, the depth of the second execution time statistics table is set to the maximum number of instructions that the control module can distribute.
[0015] In some alternative embodiments of this example, the first counter and the second counter employ a fixed-frequency clock domain.
[0016] In some alternative embodiments of this example, the performance analysis unit is further configured as follows: In response to the performance analysis enable signal being high, it is determined that the performance analysis function of the working unit is enabled; in response to the performance analysis enable signal being low, it is determined that the performance analysis function of the working unit is not enabled. In response to determining that the performance analysis function of the working unit is enabled, counting begins based on the first counter, wherein the count of the first counter is incremented by 1 every clock cycle. In response to determining that the performance analysis function of the working unit is not enabled, the count of the first counter is initialized to 0, and the current address to be written and the current address to be read are initialized to 0.
[0017] Secondly, embodiments of this application provide an instruction performance analysis method, which is applied to a reconfigurable processor. The reconfigurable processor includes a control module, multiple parallel working units, and an on-chip memory module. The instruction performance analysis method includes: Before the control module distributes instructions to the working unit, a globally unique performance analysis ID is assigned to each instruction. Based on the performance analysis enable signal sent by the control module, determine whether the performance analysis function of the working unit is enabled; In response to the determination that the performance analysis function of the working unit is enabled, based on the instruction information of each instruction sent by the working unit and the instruction execution mode of the working unit, the corresponding instruction performance analysis operation is performed on the working unit in a fixed clock domain to obtain the performance analysis result of the working unit, wherein the instruction information includes the performance analysis ID; Based on a polling strategy, the performance analysis results are read from the cache queue corresponding to each work unit, and the performance analysis results are exported to the on-chip storage module through the on-chip bus interface.
[0018] In some optional embodiments of this example, the working unit includes m working sub-units. Based on the instruction information of each instruction sent by the working unit and the instruction execution mode of the working unit, a corresponding instruction performance analysis operation is performed on the working unit in a fixed clock domain to obtain the performance analysis result of the working unit, including: In response to the instruction execution mode being a sequential execution mode, based on the instruction information sent by the working unit, a sequential instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information further includes an instruction start signal and an instruction end signal; In response to the instruction execution mode being out-of-order execution mode, based on the instruction information of each instruction sent by the working unit, an out-of-order instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information also includes the instruction's initial ID, instruction start signal, and instruction end signal.
[0019] In some optional embodiments of this example, the steps of the sequential instruction performance analysis sub-operation include: In response to receiving the instruction start signal, the current address to be written is determined based on the previous write address, and the instruction start time is determined based on a pre-set first counter; Based on the current address to be written, the performance analysis ID, and the instruction start time, update the first execution time statistics table corresponding to the working subunit; In response to receiving the instruction end signal, the current address to be read is determined based on the previous read address; Based on the current address to be read, the first execution time statistics table is read, and the read valid output data is sent to the cache queue corresponding to the work unit.
[0020] In some optional embodiments of this example, the steps of the out-of-order instruction performance analysis sub-operation include: In response to receiving the instruction start signal, the current address to be written is determined based on the initial ID of the instruction corresponding to the instruction start signal, and the instruction start time is determined based on a pre-set first counter; Based on the current address to be written, the performance analysis ID, and the instruction start time, update the second execution time statistics table corresponding to the working subunit; In response to receiving the instruction end signal, the current address to be read is determined based on the initial ID of the instruction corresponding to the instruction end signal; Based on the current address to be read, the second execution time statistics table is read, and the read valid output data is sent to the cache queue corresponding to the work unit.
[0021] In some optional embodiments of this example, determining whether the performance analysis function of the working unit is enabled based on the performance analysis enable signal sent by the control module includes: In response to the performance analysis enable signal being high, it is determined that the performance analysis function of the working unit is enabled; in response to the performance analysis enable signal being low, it is determined that the performance analysis function of the working unit is not enabled. The instruction performance analysis method also includes: In response to determining that the performance analysis function of the working unit is enabled, counting begins based on the first counter, wherein the count of the first counter is incremented by 1 every clock cycle. In response to determining that the performance analysis function of the working unit is not enabled, the count of the first counter is initialized to 0, and the current address to be written and the current address to be read are initialized to 0.
[0022] Thirdly, embodiments of this application provide a chip including the reconfigurable processor system described in any one of the first aspects.
[0023] Fourthly, embodiments of this application provide a board card including the chip described in the third aspect.
[0024] Fifthly, embodiments of this application provide an electronic device including the board described in the fourth aspect.
[0025] The present invention provides a reconfigurable processor system and instruction performance analysis method. All time statistics are completed in an independent and fixed-frequency dedicated clock domain, thereby isolating the influence of the processor core's variable clock domain. The control module assigns a globally unique performance analysis ID to each distributed instruction, enabling unambiguous tracking and association throughout the system. This allows for accurate statistics, unified management, and efficient reporting of instruction execution time for each working unit within the reconfigurable processor without significantly increasing system area and power consumption. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 This is a schematic diagram of the reconfigurable processor system provided in an embodiment of this application; Figure 2 This is one of the architectural diagrams of the performance statistics unit provided in the embodiments of this application; Figure 3 A schematic diagram of the first execution time statistics table provided in the embodiments of this application; Figure 4 A second schematic diagram of the architecture of the performance statistics unit provided in the embodiments of this application; Figure 5 A schematic diagram of the second execution time statistics table provided in the embodiments of this application; Figure 6 A schematic diagram of the architecture of the data export module provided in an embodiment of this application; Figure 7 One of the schematic diagrams of the operating clock domain for performance statistics provided in the embodiments of this application; Figure 8 A second schematic diagram of the operating clock domain for performance statistics provided in this application embodiment; Figure 9 One of the flowcharts illustrating the instruction performance analysis method provided in this application embodiment; Figure 10 A second schematic flowchart illustrating the instruction performance analysis method provided in this application embodiment; Figure 11 The third flowchart illustrating the instruction performance analysis method provided in this application embodiment; Figure 12 The fourth flowchart illustrating the instruction performance analysis method provided in this application embodiment; Figure 13 Fifth flowchart illustrating the instruction performance analysis method provided in this application embodiment; Figure 14 A flowchart illustrating the instruction performance analysis method provided in this application embodiment is shown in Figure 6. Figure 15 A block diagram of an electronic device used to implement instruction performance analysis methods is shown. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0028] The information collected in the technical solution of this application is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.
[0029] The acquisition, transmission, storage, use, and processing of data in this application comply with relevant national laws and regulations. It should be noted that certain software, components, models, and other existing industry solutions may be mentioned in the embodiments of this application. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0030] Analyzing the instruction performance of reconfigurable processors can help chip designers, compiler developers, and others identify performance bottlenecks in the working units of a reconfigurable processor, enabling targeted optimization. Existing solutions typically rely on theoretical estimations, software simulations, or rough statistics of the working state of the working units over a period of time, which has the following problems: it is impossible to guarantee the accuracy of time statistics under variable clock domains; it is difficult to accurately obtain the actual execution time of each instruction; it is difficult to support accurate instruction tracing in sequential and out-of-order execution environments at low cost; and due to the large number of instructions, performance data often consumes a lot of hardware resources and has low analysis efficiency.
[0031] In view of this, this application proposes a reconfigurable processor system, such as Figure 1 As shown, it includes a reconfigurable processor, a data export module, and a performance analysis module. The reconfigurable processor includes a control module, multiple parallel working units, and an on-chip storage module, wherein: Multiple parallel working units are designated as working unit 1, working unit 2, ..., working unit N, where N is a natural number greater than 0. The control module is configured to generate a performance analysis enable signal (i.e., profiling enable signal) and assign a globally unique corresponding performance analysis ID (i.e., profiling cmd ID) to each instruction before distributing the instruction to the working unit.
[0032] It should be understood that the control module is used to generate work instructions and distribute them to each work unit. In this embodiment, in order to accurately track each instruction, the control module assigns a globally unique performance analysis ID, i.e., Profiling cmd ID, to each instruction before distributing the instructions to the work unit. This performance analysis ID corresponds one-to-one with the instruction start time (Run time) and instruction execution time (Start time), and is exported to the on-chip storage module as the performance analysis result. In other words, the performance analysis ID is used to uniquely identify each instruction to be analyzed, ensuring that the analysis results can be traced back to the specific instruction stream.
[0033] The performance analysis module includes performance analysis units corresponding to each of the working units, i.e., the performance analysis module includes performance analysis unit 1, performance analysis unit 2, ..., performance analysis unit N; for example... Figure 1 As shown, each work unit needs to instantiate a performance analysis unit (i.e., profiling_runtime_top). The performance analysis unit is responsible for statistically analyzing the instruction start time and instruction execution time of each instruction in the work unit. Each performance analysis unit is configured as follows: Based on the performance analysis enable signal sent by the control module, it is determined whether the performance analysis function of the working unit is enabled; in response to determining that the performance analysis function of the working unit is enabled, based on the instruction information of each instruction sent by the working unit and the instruction execution mode of the working unit, the corresponding instruction performance analysis operation is performed on the working unit in a fixed clock domain to obtain the performance analysis result of the working unit, wherein the instruction information includes the performance analysis ID.
[0034] The data export module is configured to: read the performance analysis results from the cache queue corresponding to each working unit based on a polling strategy, and export the performance analysis results to the on-chip storage module through the on-chip bus interface.
[0035] In this embodiment, the control module notifies each performance analysis unit to start profiling statistics (i.e., performance analysis statistics) by sending a profiling enable signal (i.e., performance analysis enable signal). After each performance analysis unit completes the statistics, it reports the data to the upload module (i.e., the aforementioned data export module). The upload module converts the data into an AXI standard bus and submits it to the on-chip NoC (i.e., the aforementioned on-chip storage module), and then the NoC writes it to the designated storage.
[0036] It should be noted that the performance analysis unit maintains a time_count counter, which is the first counter. When the performance analysis enable signal (Profiling_enable) is high (1), the performance analysis function is enabled, and the first counter starts counting, incrementing by 1 every clock cycle; when the performance analysis enable signal is low (0), the performance analysis function is disabled, and the first counter remains at 0.
[0037] The performance analysis unit contains n performance analysis subunits (i.e., Profiling_runtime_units), the number of which is configurable and the same as the number of working subunits contained in the working unit. For example, if the working unit includes m working subunits and the performance analysis unit includes n performance analysis subunits, then n equals m, where n is a natural number greater than 0. Each performance analysis subunit maintains a set of execution time statistics tables (i.e., runtime tables).
[0038] When a command start signal is received, the performance analysis unit sends the command information issued by the working unit, such as the command start signal (i.e., Start identifier), the performance analysis ID (i.e., Profiling cmd ID), and the current time_count (which is used as the command start time), to the corresponding performance subunit (i.e., Profiling_runtime_unit) for command registration. After registration, the performance analysis subunit starts timing the execution time of the command.
[0039] When the instruction end signal is received, the performance analysis unit sends the instruction end signal (i.e., the Cmd_done identifier) issued by the working unit to the corresponding performance analysis subunit. After receiving the instruction done identifier, the performance analysis subunit stops timing and outputs relevant data such as the performance analysis ID (i.e., Profiling cmd ID), instruction start time (i.e., Start time), and instruction execution time (i.e., Run time) corresponding to the current instruction.
[0040] It should be noted that when a performance analysis unit contains multiple performance analysis sub-units, the performance analysis unit sequentially collects valid output data from each performance analysis sub-unit using a round-robin (RR) strategy and sends it to the data export module. To avoid backpressure on the main working unit, data that is not polled will be discarded. When a performance analysis unit contains only one performance analysis sub-unit, there is no RR scheduling logic.
[0041] As mentioned above, multiple instructions can be executed in parallel within a working unit, and their completion order may be sequential or out of order. In this application, in order to accurately obtain the execution time of each instruction in both sequential and out-of-order execution environments, each of the performance analysis subunits is configured as follows: In response to the instruction execution mode being a sequential execution mode, based on the instruction information sent by the working unit, a sequential instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information further includes an instruction start signal and an instruction end signal.
[0042] In response to the instruction execution mode being out-of-order execution mode, based on the instruction information of each instruction sent by the working unit, an out-of-order instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information also includes the instruction's initial ID, instruction start signal, and instruction end signal.
[0043] In this embodiment, the scheme of sequential instruction performance analysis sub-operation is suitable for work units that achieve instruction pipeline, that is, work units that start executing instructions sequentially and complete them sequentially.
[0044] In some optional embodiments of this example, in response to the instruction execution mode being a sequential execution mode, the architecture of the performance analysis subunit is as follows: Figure 2 As shown, according to Figure 2 It can be seen that each performance analysis subunit maintains a corresponding execution time statistics table. Figure 2 In this context, Time_cnt refers to the aforementioned time count counter; Profiling_cmd_start refers to the aforementioned instruction start signal; Profiling_cmd_id refers to the aforementioned performance analysis ID; and Cmd_done refers to the aforementioned instruction end signal. The performance analysis subunit is further configured as follows: In response to receiving the instruction start signal, the current address to be written is determined based on the previous write address, and the instruction start time is determined based on a pre-set first counter; Based on the current address to be written, the performance analysis ID, and the instruction start time, update the first execution time statistics table corresponding to the working subunit; In response to receiving the instruction end signal, the current address to be read is determined based on the previous read address; Based on the current address to be read, the first execution time statistics table is read, and the read valid output data is sent to the cache queue corresponding to the work unit.
[0045] It should be noted that the first execution time statistics table includes an ID field, a start time field, an execution time field, and an instruction status field. The performance analysis subunit is further configured as follows: Based on the current address to be written, write the performance analysis ID into the number field of the first execution time statistics table, write the instruction start time into the start time field, and write 1 into the instruction status field; In response to the instruction status field being set to 1, a second counter is started to keep time until the instruction end signal is received, the instruction execution time is obtained, and the instruction execution time is written into the execution time field.
[0046] For example, such as Figure 3 As shown, the first execution time statistics table includes a number field (i.e., Profiling_cmd_id), a start time field (i.e., Start_time), an execution time field (i.e., Run_time), and an instruction status field (i.e., Running). The descriptions of each field are as follows: `Profiling_cmd_id`: Represents the currently stored instruction performance analysis ID. This ID is unique and consecutive in Profiling mode (i.e., performance analysis mode) and is assigned after the Profiling function is enabled. This field is output when the `Cmd_done` signal is pulled high.
[0047] Start_time: Records the start time of the instruction, provided by the first counter, time_count. This field is output when the Cmd_done signal is pulled high.
[0048] Running: Used for status flags. It is set to 1 when Profiling_cmd_start=1 (i.e., the Profiling_cmd_start signal is triggered to a high level); and cleared to 0 when Cmd_done=1 (i.e., the Cmd_done signal is pulled high).
[0049] Run_time: Records the instruction execution time, incrementing every cycle when Running=1. This field is output when the Cmd_done signal is pulled high.
[0050] The following section will introduce how to update and read each field: Write address maintenance: The write address is write_addr. When the performance analysis enable signal is low (0), write_addr is initialized and cleared to zero, that is, the address is reset (returning to address 0). When the performance analysis enable signal is high (1), write_addr increments with each new instruction written, and after reaching the maximum address, it wraps around (i.e., writes in a loop) and returns to address 0.
[0051] Read address maintenance: The read address is read_addr: When the performance analysis enable signal is low (0), read_addr is initialized and cleared to zero, that is, the address is reset (returning to address 0); when the performance analysis enable signal is high (1) and the Cmd_done signal is triggered (that is, the Cmd_done signal is pulled high), read_addr increments by 1, wraps around (circularly writes) after reaching the maximum address, and returns 0.
[0052] New instruction writing: When the instruction start signal (i.e., the Profiling_cmd_start signal) is triggered (i.e., Profiling_cmd_start=1), new instruction information will be written to the location pointed to by write_addr. Among them, the Start_time field is written to the current time_count of the first counter; the Running field is set to 1; and Run_time starts to accumulate, that is, it starts counting based on the pre-set second counter until the instruction end signal is received, so as to obtain the instruction execution time and write the instruction execution time into the Run_time field.
[0053] Instruction result reading: When the Cmd_done signal is pulled high, the content pointed to by read_addr in the first execution time statistics table will be read out and sent to the output port of the performance analysis subunit; at the same time, Profiling_runtime_valid is pulled high, indicating that the data is valid.
[0054] It should be noted that the minimum depth of the first execution time statistics table mentioned above is set to the maximum number of instruction pipelines; the maximum depth is set to the maximum number of instructions that the control module can distribute. For example... Figure 3 As shown, if the maximum number of instruction pipelines is 4, the depth of the first execution time statistics table is 4, and the addresses are represented by labels 0-3.
[0055] In other words, the depth of the first execution time statistics table only needs to match the maximum throughput of the instruction pipeline. Therefore, the hardware area is small and no instructions are lost. It can achieve accurate statistics, unified management and efficient reporting of instruction execution time of each working unit in the reconfigurable processor without significantly increasing the system area and power consumption.
[0056] In this embodiment, the out-of-order instruction performance analysis sub-operation scheme is applicable to work units where the instruction completion order is inconsistent with the start order, and performance analysis data cannot be read according to the instruction start order.
[0057] In response to the instruction execution order being out-of-order, the architecture of the performance analysis subunit is as follows: Figure 4 As shown, Figure 4 In this context, Time_cnt refers to the aforementioned time count counter; Profiling_cmd_start refers to the aforementioned instruction start signal; Profiling_cmd_id refers to the aforementioned performance analysis ID; Cmd_done refers to the aforementioned instruction end signal; cmd_id_done is the initial ID corresponding to the instruction end signal; cmd_id_start is the initial ID corresponding to the instruction start signal; the performance analysis subunit is further configured as follows: In response to receiving the instruction start signal, the current address to be written is determined based on the initial ID of the instruction corresponding to the instruction start signal, and the instruction start time is determined based on a pre-set first counter; Based on the current address to be written, the performance analysis ID, and the instruction start time, update the second execution time statistics table corresponding to the working subunit; In response to receiving the instruction end signal, the current address to be read is determined based on the initial ID of the instruction corresponding to the instruction end signal; Based on the current address to be read, the second execution time statistics table is read, and the read valid output data is sent to the cache queue corresponding to the work unit.
[0058] The second execution time statistics table includes an ID field, a start time field, an execution time field, and an instruction status field. Updating the second execution time statistics table corresponding to the work unit based on the current address to be written, the performance analysis ID, and the instruction start time includes: Based on the current address to be written, write the performance analysis ID into the number field of the second execution time statistics table, write the instruction start time into the start time field, and write 1 into the instruction status field; In response to the instruction status field being set to 1, a second counter is started to keep time until the instruction end signal is received, the instruction execution time is obtained, and the instruction execution time is written into the execution time field.
[0059] For example, such as Figure 5As shown, the second execution time statistics table includes a number field (i.e., Profiling_cmd_id), a start time field (i.e., Start_time), an execution time field (i.e., Run_time), and an instruction status field (i.e., Running). The descriptions of each field are as follows: `Profiling_cmd_id`: Represents the currently stored instruction performance analysis ID. This ID is unique and consecutive in Profiling mode and is assigned after the Profiling function is enabled. This field is output when the `Cmd_done` signal is pulled high.
[0060] Start_time: Records the start time of the instruction, provided by the first counter, time_count. This field is output when the Cmd_done signal is pulled high.
[0061] Running: Used for status flags. It is set to 1 when Profiling_cmd_start=1 (i.e., the Profiling_cmd_start signal is triggered to a high level); and cleared to 0 when Cmd_done=1 (i.e., the Cmd_done signal is pulled high).
[0062] Run_time: Records the instruction execution time, incrementing every cycle when Running=1. This field is output when Cmd_done goes high.
[0063] The following section will introduce how to update and read each field: Write address maintenance: The write address is write_addr. When Profiling_enable = 0 (i.e., the performance analysis enable signal is low), write_addr is initialized and cleared to zero, that is, the address is reset. When Profiling_enable = 1 (i.e., the performance analysis enable signal is high level 1), write_addr uses the initial ID corresponding to the instruction start signal as the address.
[0064] Read address maintenance: The read address is read_addr: When the performance analysis enable signal is low (0), read_addr is initialized and cleared; when the performance analysis enable signal is high (1) and the Cmd_done signal is triggered, read_addr uses the initial ID corresponding to the instruction end signal as the address.
[0065] New instruction writing: When the Profiling_cmd_start signal is triggered, new instruction information will be written to the location pointed to by write_addr. The Start_time field is written to the current time_count; the Running field is set to 1; and Run_time starts to accumulate, that is, it starts counting based on the pre-set second counter until the instruction end signal is received, the instruction execution time is obtained, and the instruction execution time is written to the Run_time field.
[0066] Instruction result reading: When the Cmd_done signal goes high, the content at the location pointed to by read_addr in the runtime table will be read and sent to the output port of the performance analysis subunit; at the same time, Profiling_runtime_valid goes high, indicating that the data is valid.
[0067] It should be noted that the depth of the second execution time statistics table mentioned above should be fixed to the maximum number of instructions that the control module can distribute.
[0068] For example, the control module can distribute a maximum of 32 instructions, such as... Figure 5 As shown, the depth of the second execution time statistics table is fixed at 32, and the addresses are represented by labels 0-31. For example, there are 3 instructions in a working sub-unit, and the execution order of the 3 instructions is out of order. The 3 instructions have performance analysis IDs and initial IDs, such as initial IDs 5, 10 and 12. When an instruction starts, that is, after receiving the instruction start signal, the current address to be written is determined according to the initial ID corresponding to the instruction start signal. For example, when the instruction with initial ID 5 starts execution, the current address to be written is 5. When an instruction ends, that is, after receiving the instruction end signal, the current address to be read is determined according to the initial ID corresponding to the instruction end signal. For example, when the instruction with initial ID 12 ends execution, the current address to be read is 12.
[0069] It should also be noted that the first and second counters use a fixed-frequency clock domain.
[0070] In other words, in this embodiment, the timing statistics of instruction execution are all completed in a dedicated clock domain with an independent and fixed frequency, thereby isolating the influence of the processor core's variable clock domain and fundamentally ensuring the uniformity of timing measurement benchmarks and the accuracy of statistical results.
[0071] In some optional embodiments of this example, the performance analysis unit is further configured to: determine that the performance analysis function of the working unit is enabled in response to the performance analysis enable signal being high; and determine that the performance analysis function of the working unit is not enabled in response to the performance analysis enable signal being low. In response to determining that the performance analysis function of the working unit is enabled, counting begins based on the first counter, wherein the count of the first counter is incremented by 1 every clock cycle. In response to determining that the performance analysis function of the working unit is not enabled, the count of the first counter is initialized to 0, and the current address to be written and the current address to be read are initialized to 0.
[0072] In this embodiment, after each working unit completes the execution of an instruction, it sends the relevant data of the instruction execution, including the performance analysis ID (Profiling cmd ID), the instruction start time (Start time), and the instruction execution time (Runtime), to the data export module.
[0073] The architecture of the data export module is as follows: Figure 6 As shown, the performance analysis data for each work unit (i.e., the valid data output by the performance analysis unit) is sent to an independent cache queue. The queue depth is configurable; if the queue is full, newly arriving data will be discarded. For example, as shown... Figure 6 As shown, the performance analysis data of working unit 1 (i.e., the effective output data of each performance analysis unit) is output from performance analysis unit 1 to cache queue 1, the performance analysis data of working unit 2 is output from performance analysis unit 2 to cache queue 2, that is, the performance analysis data of working unit n is output from performance analysis unit N to cache queue N.
[0074] The data export module uses a polling mechanism to read ready data from each cache queue and send it to the on-chip storage module via AXI128 (the bus width can be modified according to project requirements).
[0075] The To AXI submodule is responsible for packaging data into AXI 128 format and sending it out. The module supports writing data from different work units to different address spaces; in this case, the software needs to configure an independent AXI write address for each work unit.
[0076] It should be noted that, as Figure 7 As shown, performance statistics require precise measurement of the absolute time of instruction execution, therefore its operating clock must use a fixed reference clock, i.e., a fixed clock domain. However, the reconfigurable processor uses a variable frequency clock for information such as instruction sending by the control module, performance analysis enable signals, instruction execution by the computing unit, instruction start and end, and reporting.
[0077] The profiling enable switch (i.e., the performance analysis enable signal Profiling_enable) and instruction execution related signals (including instruction start / end flags, performance analysis ID, and initial ID) all originate from the variable clock domain. These signals must undergo cross-clock domain (CDC) processing before being connected to the performance analysis subunit. Similarly, when performing a report, since the report interface is located in the variable frequency domain while the performance analysis subunit uses a fixed frequency clock domain, the reported data also needs to undergo corresponding cross-clock domain (CDC) processing.
[0078] In other words, such as Figure 8 As shown, the input signals—Profiling Start 0 ~ n, Profiling cmd id0 ~ n, cmd done 0 ~ n, etc.—and the output signals—Profiling data valid, Profiling cmd ID, Starttime, and Run time—operate within the variable-frequency processor main clock, while other logic operates within the fixed-frequency Profiling clock. Therefore, the performance analysis unit requires two cross-clock domain processing steps.
[0079] Thus, the reconfigurable processor system provided by this invention completes all time statistics in an independent and fixed-frequency dedicated clock domain, thereby isolating the influence of the processor core's variable clock domain and fundamentally ensuring the uniformity of timing measurement benchmarks and the accuracy of statistical results. Each instruction distributed by the control module is assigned a globally unique performance analysis ID, enabling unambiguous tracking and association throughout the system. Each working unit has an independent performance analysis unit that accurately records the start time and execution duration of the executed instructions, ensuring correct binding of statistical data to the instructions themselves even in scenarios of parallel and out-of-order instruction execution. The statistical data from each working unit are collected through efficient internal pathways and ultimately written to a designated external storage area in an orderly and efficient manner via a standard on-chip bus (such as AXI) interface, providing performance profiling data for the software layer. Ultimately, without significantly increasing system area and power consumption, it achieves accurate statistics, unified management, and efficient reporting of instruction execution time for each working unit within the reconfigurable processor.
[0080] This invention also provides an instruction performance analysis method applied to a reconfigurable processor, as described in the following embodiments. Since the principle underlying this method is similar to that of a reconfigurable processor system, its implementation can be found in the implementation of a reconfigurable processor system, and repetitions will not be repeated.
[0081] like Figure 9As shown, the instruction performance analysis methods applied to reconfigurable processors include: S10. Before the control module distributes the instructions to the working unit, assign a globally unique performance analysis ID to each instruction. S20. Based on the performance analysis enable signal sent by the control module, determine whether the performance analysis function of the working unit is enabled. S30. In response to determining that the performance analysis function of the working unit is enabled, based on the instruction information of each instruction sent by the working unit and the instruction execution mode of the working unit, perform the corresponding instruction performance analysis operation on the working unit in a fixed clock domain to obtain the performance analysis result of the working unit, wherein the instruction information includes the performance analysis ID; S40. Based on a polling strategy, read the performance analysis results from the cache queue corresponding to each work unit, and export the performance analysis results to the on-chip storage module through the on-chip bus interface.
[0082] This invention provides an instruction performance analysis method for reconfigurable processors. All time statistics are performed in a dedicated clock domain with an independent and fixed frequency, thereby isolating the influence of the processor core's variable clock domain. Each instruction distributed by the control module is assigned a globally unique performance analysis ID, which enables unambiguous tracking and association throughout the system. This method achieves accurate statistics, unified management, and efficient reporting of instruction execution time for each working unit within the reconfigurable processor without significantly increasing system area and power consumption.
[0083] The following is a detailed explanation of S10-S40: S10. Before the control module distributes the instructions to the working unit, assign a globally unique performance analysis ID to each instruction.
[0084] It should be understood that the reconfigurable processor includes a control module, multiple parallel working units, and an on-chip storage module. The control module is used to generate working instructions and distribute them to each working unit. In this embodiment, in order to accurately track each instruction, the control module assigns a globally unique performance analysis ID, i.e., a Profiling cmd ID, to each instruction before distributing it to the working unit. This performance analysis ID corresponds one-to-one with the instruction start time and instruction execution time, and is exported to the on-chip storage module as the performance analysis result. In other words, the performance analysis ID is used to uniquely identify each instruction to be analyzed, ensuring that the analysis results can be traced back to the specific instruction stream.
[0085] S20. Based on the performance analysis enable signal sent by the control module, determine whether the performance analysis function of the working unit is enabled.
[0086] In this application, the control module notifies each module whether to start profiling statistics by sending a performance analysis enable signal, i.e., a profiling enable signal, where: When the performance analysis enable signal is high, it is determined that the performance analysis function of the working unit is enabled; when the performance analysis enable signal is low, it is determined that the performance analysis function of the working unit is not enabled.
[0087] It should be noted that when it is determined that the performance analysis function of the working unit is enabled, counting begins based on the first counter, wherein the count of the first counter is incremented by 1 every clock cycle; when it is determined that the performance analysis function of the working unit is not enabled, the count of the first counter is initialized to 0, and the current address to be written and the current address to be read are also initialized to 0.
[0088] S30. In response to determining that the performance analysis function of the working unit is enabled, based on the instruction information of each instruction sent by the working unit and the instruction execution mode of the working unit, perform the corresponding instruction performance analysis operation on the working unit in a fixed clock domain to obtain the performance analysis result of the working unit, wherein the instruction information includes the performance analysis ID.
[0089] As mentioned above, multiple instructions can be executed in parallel within a working unit, and their completion order may be sequential or out of order. In this application, in order to accurately obtain the execution time of each instruction in both sequential and out-of-order execution environments, the working unit includes m working sub-units, such as... Figure 10 As shown, S30 further includes: S301. In response to the instruction execution mode being a sequential execution mode, based on the instruction information of each instruction sent by the working unit, a sequential instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information further includes an instruction start signal and an instruction end signal. S302. In response to the instruction execution mode being out-of-order execution mode, based on the instruction information of each instruction sent by the working unit, an out-of-order instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information also includes the instruction's initial ID, instruction start signal, and instruction end signal.
[0090] In some alternative embodiments of this example, the sequential instruction performance analysis sub-operation scheme is suitable for working units that implement instruction pipelines, i.e., working units that start and finish instruction execution sequentially.
[0091] like Figure 11As shown, the steps of the sequential instruction performance analysis sub-operation include: S3011. In response to receiving the instruction start signal, determine the current address to be written based on the previous write address, and determine the instruction start time based on a pre-set first counter.
[0092] It should be noted that the instruction information issued by the working unit includes the performance analysis ID, the instruction start signal (i.e., the Profiling_cmd_Start signal), and the instruction end signal (Cmd_done signal). When the instruction start signal is received, the instruction start time can be determined based on the pre-set first counter.
[0093] When the aforementioned performance analysis function is enabled, the first counter starts counting and increments by 1 every clock cycle. When an instruction start signal is received, the instruction start time can be determined based on the current count of the first counter, which is set in advance.
[0094] Furthermore, for work units where instructions are executed sequentially and completed sequentially, when an instruction start signal is received, write_addr increments by 1 based on the address of the previous write operation to determine the current address to be written.
[0095] S3012. Based on the current address to be written, the performance analysis ID, and the instruction start time, update the first execution time statistics table corresponding to the working subunit.
[0096] like Figure 12 As shown, updating the first execution time statistics table based on the current address to be written, the performance analysis ID, and the instruction start time includes: S30121. Based on the current address to be written, write the performance analysis ID into the number field of the first execution time statistics table, write the instruction start time into the start time field, and write 1 into the instruction status field.
[0097] S30122. In response to the instruction status field being set to 1, a second counter is started to keep time until the instruction end signal is received, the instruction execution time is obtained, and the instruction execution time is written into the execution time field.
[0098] S3013. In response to receiving the instruction end signal, determine the current address to be read based on the previous read address.
[0099] In this step, when Profiling_enable = 0, read_addr is initialized to zero; when Profiling_enable = 1 and the cmd_done signal is triggered, read_addr increments by 1 based on the previous read address to determine the current address to be read.
[0100] S3014. Based on the current address to be read, read the first execution time statistics table and send the read valid output data to the cache queue corresponding to the work unit.
[0101] In this step, when the cmd_done signal goes high, the content of the entry pointed to by read_addr in the runtime table will be read out and sent to the Profiling_runtime_unit output port; at the same time, Profiling_runtime_valid goes high, indicating that the data is valid.
[0102] In some optional embodiments of this example, the out-of-order instruction performance analysis sub-operation scheme is applicable to work units where the instruction completion order is inconsistent with the start order and performance analysis data cannot be read according to the instruction start order.
[0103] like Figure 13 As shown, the steps of the out-of-order instruction performance analysis sub-operation include: S3021. In response to receiving the instruction start signal, determine the current address to be written based on the initial ID of the instruction corresponding to the instruction start signal, and determine the instruction start time based on a pre-set first counter.
[0104] In this step, unlike the sequential instruction performance analysis sub-operation mentioned above, the instruction completion order of the corresponding work unit in the out-of-order instruction performance analysis sub-operation may be inconsistent with the start order. Therefore, performance analysis data cannot be read according to the instruction start order. Thus, the cmd_id signal (i.e., the initial ID) from the interface signals of the control module and each work unit needs to be used as the unique identifier for instruction registration and submission. It should be noted that cmd_id is used to represent the instructions distributed by the control module, not the Profiling_cmd_id used for instruction performance statistics. If the control module distributes a maximum of 32 instructions simultaneously, then cmd_id[4:0] can be used.
[0105] When the instruction start signal is received, the current address to be written is determined based on the initial ID of the instruction corresponding to the instruction start signal, and the instruction start time is determined based on a pre-set first counter.
[0106] S3022. Based on the current address to be written, the performance analysis ID, and the instruction start time, update the second execution time statistics table corresponding to the working subunit.
[0107] like Figure 14 As shown, updating the second execution time statistics table corresponding to the work unit based on the current address to be written, the performance analysis ID, and the instruction start time includes: S30221. Based on the current address to be written, write the performance analysis ID into the number field of the second execution time statistics table, write the instruction start time into the start time field, and write 1 into the instruction status field. S30222. In response to the instruction status field being set to 1, a second counter is started to keep time until the instruction end signal is received, the instruction execution time is obtained, and the instruction execution time is written into the execution time field.
[0108] S3023. In response to receiving the instruction end signal, determine the current address to be read based on the initial ID of the instruction corresponding to the instruction end signal.
[0109] S3024. Based on the current address to be read, read the second execution time statistics table and send the read valid output data to the cache queue corresponding to the working unit.
[0110] In this embodiment, when an instruction begins execution, it is written to the runtime table using cmd_id as the address; when the instruction ends execution, its performance data is read from the runtime table using cmd_id as the address.
[0111] Instructions are distinguished by cmd_id. The depth of the runtime table must be fixed to the maximum number of instructions that the control module can distribute. This is suitable for scenarios that require tracking out-of-order instructions.
[0112] S40. Based on a polling strategy, read the performance analysis results from the cache queue corresponding to each work unit, and export the performance analysis results to the on-chip storage module through the on-chip bus interface.
[0113] In this embodiment, after each work unit completes the execution of an instruction, it sends the relevant data of the instruction execution, including the Profiling cmd id, Start time, and Run time, to the data export module. The performance analysis data of each work unit is sent to an independent cache queue. The queue depth is configurable; if the queue is full, newly arrived data will be discarded.
[0114] A polling mechanism is used to read ready data from each cache queue and send it to the on-chip storage module NoC via AXI128 (the bus width can be modified according to project requirements).
[0115] The To AXI submodule is responsible for packaging data into AXI128 format and sending it out. The module supports writing data from different work units to different address spaces; in this case, the software needs to configure an independent AXI write address for each work unit.
[0116] It should be noted that the depth, column field descriptions, updates and reading of the first and second execution time statistics tables mentioned above are the same as those in the aforementioned embodiments, and will not be repeated here.
[0117] It should be noted that the first and second counters operate in a fixed-frequency clock domain. A fixed-frequency clock is used for performance statistics, and communication with the main clock domain is resolved through two CDC (Continuous Damping and Control) steps, solving the problem of not being able to guarantee the accuracy of time statistics in a variable clock domain. Instruction performance statistics are performed within each working unit, accurately calculating the actual execution time of each instruction using start / stop signals within the working unit, solving the problem of difficulty in accurately obtaining the actual execution time of each instruction. Two solutions are provided: the sequential solution requires the runtime table depth to match only the maximum throughput of the instruction pipeline, resulting in a smaller hardware area; the out-of-order solution requires the runtime table depth to be fixed to the maximum number of instructions that the control module can distribute, solving the problem of not being able to support accurate instruction tracking in both sequential and out-of-order execution environments at low cost. Through centralized management and unified reporting, each working module does not need a large amount of hardware resources to record performance data; it only needs to write the performance data to a designated memory, resulting in low cost and easy software reading and analysis, solving the problem that performance data collection and analysis often consumes a lot of hardware resources and is inefficient due to the large number of instructions.
[0118] Thus, the instruction performance analysis method provided in this embodiment completes all time statistics in an independent and fixed-frequency dedicated clock domain, thereby isolating the influence of the processor core's variable clock domain. Each instruction distributed by the control module is assigned a globally unique ID for unambiguous tracking and association throughout the system. Each working unit has an independent performance analysis unit that can accurately record the start time and execution duration of the executed instructions, and ensure that the statistical data and the instructions themselves are correctly bound even in scenarios where instructions are executed in parallel or out of order. The statistical data of each working unit are collected through an efficient internal path and finally written to a designated external storage area in an orderly and efficient manner through a standard on-chip bus (such as AXI) interface, providing performance profiling data for the software layer.
[0119] This application also provides a chip that includes the reconfigurable processor system described in the foregoing embodiments.
[0120] This application also provides a board card that includes the chip described in the foregoing embodiments.
[0121] like Figure 15 As shown, this application embodiment also includes an electronic device. Figure 15 This is a block diagram illustrating an electronic device 900 for implementing the above-described instruction performance analysis method, according to an exemplary embodiment. For example, the electronic device 900 may be an AI server, a training and promotion integrated machine, etc.
[0122] Reference Figure 15 The electronic device 900 may include one or more of the following components: an AI-accelerated computing module, a CPU module, a power supply module, a hard drive module, and a fan module. Each module works in conjunction with the bus system through a standardized hardware interface, with the specific architecture as follows: The AI-accelerated computing module comprises multiple AI accelerator cards deployed in parallel. Each AI accelerator card integrates at least one AI accelerator chip (such as an RPU chip, GPU chip, or CGRA chip). Data communication between the AI accelerator cards is achieved through a high-speed card-to-card (C2C) interconnect structure, supporting low-latency, high-bandwidth horizontal scaling. The AI accelerator chip is dedicated to performing AI computing tasks such as high-density matrix operations, neural network model training, and / or inference, providing the main computing power support.
[0123] The CPU module includes at least one CPU board, which houses a central processing unit (CPU) chip and associated CPU memory (such as DDR4 / DDR5, RAM). The CPU chip serves as the system control center, responsible for task scheduling, resource allocation, I / O management, and coordinating the parallel computing of the AI acceleration computing module, while also handling non-accelerated general-purpose computing tasks.
[0124] The power module is equipped with redundant power supply units to provide stable power distribution and management for the AI acceleration computing module, CPU module and other modules.
[0125] The hard drive module integrates a high-speed solid-state drive (SSD) and / or a large-capacity hard disk drive (HDD), connected to the system bus via a backplane. The hard drive stores the operating system, AI training datasets, model parameters, and computation results, providing high-throughput data read / write channels and supporting data preprocessing and persistence.
[0126] The fan module uses a multi-zone independent speed-controlled fan array, which is configured in key heat source areas (such as AI accelerator cards and CPU heat dissipation areas) to achieve system heat dissipation through forced air cooling and ensure the stable operation of high-efficiency computing components.
[0127] The CPU module is connected to the AI acceleration computing module via the PCIe bus to enable task distribution, result collection, and memory coordination.
[0128] The CPU module manages the data access of the hard drive module through SATA / SAS / NVMe interfaces.
[0129] The power module provides tiered power to all functional modules through the power distribution backplane.
[0130] The fan module adjusts the fan speed based on temperature monitoring signals from the CPU board and AI accelerator card.
[0131] It should be noted that, in this application, apart from the different meanings of the number of working units N and the number of working subunits n, the capitalization of other letters is not used to distinguish meanings. That is, the same letter or letter sequence, whether uppercase or lowercase, represents the same technical meaning. For example, "id" and "ID", "cmd" and "Cmd", "Start" and "start", and "Profiling" and "profiling" should all be understood as referring to the same object or parameter, unless the context of this application explicitly indicates otherwise. Any interpretation that asserts different meanings based on capitalization differences is not considered an intention of this application.
[0132] It should be noted that in the description of this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0133] In the embodiments of this application, the singular forms "a," "the," etc., including the plural forms, should be broadly understood as "a kind" or "a class" rather than limited to the meaning of "an." Furthermore, the term "the" should be understood to include both the singular and plural forms, unless the context explicitly indicates otherwise. Additionally, the term "according to" should be understood as "at least partially based on…," and the term "based on" should be understood as "at least partially based on…," unless the context explicitly indicates otherwise.
[0134] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0135] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A reconfigurable processor system, characterized in that, It includes a reconfigurable processor, a data export module, and a performance analysis module. The reconfigurable processor includes a control module, multiple parallel working units, and an on-chip storage module, wherein: The control module is configured to generate a performance analysis enable signal and assign a globally unique performance analysis ID to each instruction before distributing the instructions to the work unit. The performance analysis module includes performance analysis units corresponding to each of the working units, wherein each of the performance analysis units is configured as follows: Based on the performance analysis enable signal sent by the control module, it is determined whether the performance analysis function of the working unit is enabled; in response to determining that the performance analysis function of the working unit is enabled, based on the instruction information of each instruction sent by the working unit and the instruction execution mode of the working unit, the corresponding instruction performance analysis operation is performed on the working unit in a fixed clock domain to obtain the performance analysis result of the working unit, wherein the instruction information includes the performance analysis ID; The data export module is configured to: read the performance analysis results from the cache queue corresponding to each working unit based on a polling strategy, and export the performance analysis results to the on-chip storage module through the on-chip bus interface.
2. The reconfigurable processor system according to claim 1, wherein the working unit comprises m working sub-units, characterized in that, The performance analysis unit comprises n performance analysis subunits, where n equals m and n is a natural number greater than 0. Each performance analysis subunit is configured as follows: In response to the instruction execution mode being a sequential execution mode, based on the instruction information sent by the working unit, a sequential instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information further includes an instruction start signal and an instruction end signal; In response to the instruction execution mode being out-of-order execution mode, based on the instruction information of each instruction sent by the working unit, an out-of-order instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information also includes the instruction's initial ID, instruction start signal, and instruction end signal.
3. The reconfigurable processor system according to claim 2, characterized in that, When n is greater than 1, the performance analysis unit is further configured as follows: The effective output data of each performance analysis subunit is collected sequentially using a polling strategy, and the effective output data is sent to the cache queue corresponding to the working unit.
4. The reconfigurable processor system according to claim 2, characterized in that, In response to the instruction execution mode being a sequential execution mode, the performance analysis subunit is further configured as follows: In response to receiving the instruction start signal, the current address to be written is determined based on the previous write address, and the instruction start time is determined based on a pre-set first counter; Based on the current address to be written, the performance analysis ID, and the instruction start time, update the first execution time statistics table corresponding to the working subunit; In response to receiving the instruction end signal, the current address to be read is determined based on the previous read address; Based on the current address to be read, the first execution time statistics table is read, and the read valid output data is sent to the cache queue corresponding to the work unit.
5. The reconfigurable processor system according to claim 4, characterized in that, The first execution time statistics table includes an ID field, a start time field, an execution time field, and an instruction status field. The performance analysis subunit is further configured as follows: Based on the current address to be written, write the performance analysis ID into the number field of the first execution time statistics table, write the instruction start time into the start time field, and write 1 into the instruction status field; In response to the instruction status field being set to 1, a second counter is started to keep time until the instruction end signal is received, the instruction execution time is obtained, and the instruction execution time is written into the execution time field.
6. The reconfigurable processor system according to claim 5, characterized in that, The minimum depth of the first execution time statistics table is set to the maximum number of instruction pipelines; the maximum depth of the first execution time statistics table is set to the maximum number of instructions that the control module can distribute.
7. The reconfigurable processor system according to claim 2, characterized in that, In response to the instruction execution order being out-of-order execution mode, the performance analysis subunit is further configured as follows: In response to receiving the instruction start signal, the current address to be written is determined based on the initial ID of the instruction corresponding to the instruction start signal, and the instruction start time is determined based on a pre-set first counter; Based on the current address to be written, the performance analysis ID, and the instruction start time, update the second execution time statistics table corresponding to the working subunit; In response to receiving the instruction end signal, the current address to be read is determined based on the initial ID of the instruction corresponding to the instruction end signal; Based on the current address to be read, the second execution time statistics table is read, and the read valid output data is sent to the cache queue corresponding to the work unit.
8. The reconfigurable processor system according to claim 7, characterized in that, The second execution time statistics table includes an ID field, a start time field, an execution time field, and an instruction status field. Updating the second execution time statistics table corresponding to the work unit based on the current address to be written, the performance analysis ID, and the instruction start time includes: Based on the current address to be written, write the performance analysis ID into the number field of the second execution time statistics table, write the instruction start time into the start time field, and write 1 into the instruction status field; In response to the instruction status field being set to 1, a second counter is started to keep time until the instruction end signal is received, the instruction execution time is obtained, and the instruction execution time is written into the execution time field.
9. The reconfigurable processor system according to claim 8, characterized in that, The depth of the second execution time statistics table is set to the maximum number of instructions that the control module can distribute.
10. The reconfigurable processor system according to any one of claims 4-9, characterized in that, The first and second counters use a fixed-frequency clock domain.
11. The reconfigurable processor system according to any one of claims 4 or 7, characterized in that, The performance analysis unit is further configured to: In response to the performance analysis enable signal being high, it is determined that the performance analysis function of the working unit is enabled; in response to the performance analysis enable signal being low, it is determined that the performance analysis function of the working unit is not enabled. In response to determining that the performance analysis function of the working unit is enabled, counting begins based on the first counter, wherein the count of the first counter is incremented by 1 every clock cycle. In response to determining that the performance analysis function of the working unit is not enabled, the count of the first counter is initialized to 0, and the current address to be written and the current address to be read are initialized to 0.
12. A method for instruction performance analysis, characterized in that, The method is applied to a reconfigurable processor, which includes a control module, multiple parallel working units, and an on-chip memory module. The instruction performance analysis method includes: Before the control module distributes instructions to the working unit, a globally unique performance analysis ID is assigned to each instruction. Based on the performance analysis enable signal sent by the control module, determine whether the performance analysis function of the working unit is enabled; In response to the determination that the performance analysis function of the working unit is enabled, based on the instruction information of each instruction sent by the working unit and the instruction execution mode of the working unit, the corresponding instruction performance analysis operation is performed on the working unit in a fixed clock domain to obtain the performance analysis result of the working unit, wherein the instruction information includes the performance analysis ID; Based on a polling strategy, the performance analysis results are read from the cache queue corresponding to each work unit, and the performance analysis results are exported to the on-chip storage module through the on-chip bus interface.
13. The instruction performance analysis method according to claim 12, wherein the working unit comprises m working sub-units, characterized in that, Based on the instruction information of each instruction sent by the working unit and the instruction execution mode of the working unit, a corresponding instruction performance analysis operation is performed on the working unit in a fixed clock domain to obtain the performance analysis results of the working unit, including: In response to the instruction execution mode being a sequential execution mode, based on the instruction information sent by the working unit, a sequential instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information further includes an instruction start signal and an instruction end signal; In response to the instruction execution mode being out-of-order execution mode, based on the instruction information of each instruction sent by the working unit, an out-of-order instruction performance analysis sub-operation is performed on the working sub-unit in a fixed clock domain, wherein the instruction information also includes the instruction's initial ID, instruction start signal, and instruction end signal.
14. The instruction performance analysis method according to claim 13, characterized in that, The steps of the sequential instruction performance analysis sub-operation include: In response to receiving the instruction start signal, the current address to be written is determined based on the previous write address, and the instruction start time is determined based on a pre-set first counter; Based on the current address to be written, the performance analysis ID, and the instruction start time, update the first execution time statistics table corresponding to the working subunit; In response to receiving the instruction end signal, the current address to be read is determined based on the previous read address; Based on the current address to be read, the first execution time statistics table is read, and the read valid output data is sent to the cache queue corresponding to the work unit.
15. The instruction performance analysis method according to claim 13, characterized in that, The steps of the out-of-order instruction performance analysis sub-operation include: In response to receiving the instruction start signal, the current address to be written is determined based on the initial ID of the instruction corresponding to the instruction start signal, and the instruction start time is determined based on a pre-set first counter; Based on the current address to be written, the performance analysis ID, and the instruction start time, update the second execution time statistics table corresponding to the working subunit; In response to receiving the instruction end signal, the current address to be read is determined based on the initial ID of the instruction corresponding to the instruction end signal; Based on the current address to be read, the second execution time statistics table is read, and the read valid output data is sent to the cache queue corresponding to the work unit.
16. The instruction performance analysis method according to claim 14 or 15, characterized in that, The determination of whether the performance analysis function of the working unit is enabled based on the performance analysis enable signal sent by the control module includes: In response to the performance analysis enable signal being high, it is determined that the performance analysis function of the working unit is enabled; in response to the performance analysis enable signal being low, it is determined that the performance analysis function of the working unit is not enabled. The instruction performance analysis method also includes: In response to determining that the performance analysis function of the working unit is enabled, counting begins based on the first counter, wherein the count of the first counter is incremented by 1 every clock cycle. In response to determining that the performance analysis function of the working unit is not enabled, the count of the first counter is initialized to 0, and the current address to be written and the current address to be read are initialized to 0.
17. A chip, characterized in that, Includes the reconfigurable processor system according to any one of claims 1-11.
18. A circuit board, characterized in that, Includes the chip described in claim 17.
19. An electronic device, characterized in that, Includes the board as described in claim 18.