A method for monitoring the execution of pipeline instructions

By obtaining data at all levels of the CPU pipeline and creating hierarchical table entries for analysis, the problem of time-consuming execution detection of CPU pipeline instructions in the prior art is solved, and rapid detection and real-time abnormal feedback are achieved.

CN115454505BActive Publication Date: 2025-07-22SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211080024.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-05
Publication Date
2025-07-22
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

The prior art consumes clock cycles in CPU pipeline instruction execution detection, resulting in lag processing when developing high-performance processors, making it difficult to quickly detect exceptions.

Method used

By obtaining data at all levels of the CPU pipeline, creating hierarchical table items for data collection and analysis, determining whether the pipeline execution is abnormal, and returning the exception signal to the main processor for processing.

Benefits of technology

It realizes rapid detection of CPU pipeline abnormalities, reduces clock cycle waste, and improves real-time performance of pipeline execution and user interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115454505B_ABST
    Figure CN115454505B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for monitoring the execution of CPU pipeline instructions, which acquires data at all levels of the CPU pipeline; creates hierarchical table entries to collect data at all levels of the CPU pipeline, where different detection items correspond to different positions in the CPU pipeline; determines whether the execution of the CPU pipeline is abnormal based on the current data at all levels of the CPU pipeline obtained from the hierarchical table entries. If the judgment result is abnormal, an abnormal signal is returned to the main processor, and the main processor makes a reaction to perform data processing according to different positions of the CPU pipeline stall or different reasons for instruction anomalies. The present invention can collect and analyze data at all levels of the pipeline, enabling the rapid detection of abnormal pipeline operation and sending it to the user, and discovering pipeline anomalies as quickly as possible. The present invention can enable the rapid detection of abnormal pipeline operation and send it to the user, and ensure the real-time detection of pipeline anomalies as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for monitoring the execution of pipeline instructions, belonging to the technical field of integrated circuit design. Background Art

[0002] Currently, the domestic R & D of mass-produced chips of high-performance CPUs based on RISC-V is facing many problems, especially there are bottlenecks in CPU performance. Therefore, the detection of CPU pipeline instruction execution and performance is particularly important.

[0003] In the existing technology, to judge the execution situation of CPU pipeline instructions, it is often necessary to collect the data of a certain module and then send the data to a specific memory management unit for comparison to judge whether the CPU pipeline execution is abnormal. And this process may consume the clock cycles of pipeline execution, which is not conducive to the R & D of high-performance processors. At the same time, this process may lead to the lag in processing when the cpu instruction execution is abnormal. Summary of the Invention

[0004] Aiming at the deficiencies of the existing technology, the present invention provides a method for monitoring the execution of cpu pipeline instructions, which can collect and analyze the data at all levels of the pipeline, so that when the pipeline operation is abnormal, it can be quickly detected and sent to the user, and the pipeline abnormality can be found as soon as possible.

[0005] The technical solution of the present invention is as follows:

[0006] A method for monitoring the execution of cpu pipeline instructions, comprising:

[0007] Obtaining the data at all levels of the cpu pipeline;

[0008] Creating hierarchical table entries, collecting the data at all levels of the cpu pipeline, and different detection items correspond to different positions of the cpu pipeline;

[0009] Judging whether the cpu pipeline execution is abnormal according to the data at all levels of the current cpu pipeline obtained from the hierarchical table entries. If the judgment result is abnormal, an abnormal signal is returned to the main processor, and the main processor makes a reaction and processes the data according to different positions of the cpu pipeline pause or different reasons for instruction abnormality.

[0010] Preferably according to the present invention, the hierarchical table entries include multiple sub-tables, and the multiple sub-tables are divided into table entries of different levels according to the number of pipeline stages in the cpu pipeline.

[0011] Preferably according to the present invention, each sub-table includes multiple detection items, and respectively processes the data at different positions of the cpu pipeline corresponding to the number of pipeline stages;

[0012] The detection items include: the detection of the front end of the CPU pipeline and the detection of the back end of the CPU pipeline; the detection of the front end of the pipeline includes the detection of the instruction fetch unit and the detection of the instruction decoding unit; the detection of the back end of the pipeline includes the detection of the instruction load / store unit and the detection of the instruction retirement unit RTU.

[0013] Preferably according to the present invention, the detection of the instruction fetch unit includes:

[0014] The instruction fetch unit IFU includes an IF1 stage, an IF2 stage, an IF3 stage, and an IF4 stage. The IF1 stage is the instruction fetch stage, the IF2 stage is the instruction renaming stage, the IF3 stage is the instruction dispatch stage, and the IF4 stage is the instruction issue stage; in the IF3 stage, it is responsible for packing and preprocessing instructions, and then sending them to the instruction decoding unit IDU for decoding various instructions; at this time, the corresponding flag signal of the instruction fetch unit IFU is transmitted to the entry index entry. The entry index entry includes entry1, entry2, entry3, and entry4, which store information collected from the pipeline; the information related to the instruction fetch unit IFU is stored in entry1; then, the main processor compares the information in the index entry with the previously set detection items to determine whether the pipeline operation in this stage is normal.

[0015] Preferably according to the present invention, the detection of the instruction decoding unit includes:

[0016] The instruction decoding unit IDU decodes four instructions simultaneously. After completing the detection of data dependencies, decomposition of micro-operations, register mapping, instruction dispatch, and instruction issue processes, it issues the instructions to the next-level CPU pipeline. At this time, the corresponding data of the instruction decoding unit IDU is transmitted to entry2, and the main processor compares the data in entry2 with the detection items to determine whether the instruction execution of the IDU-level pipeline is normal;

[0017] Specifically, it includes: checking whether the instructions entering the instruction decoding unit IDU correspond one by one to the instructions fetched and sent by the instruction fetch unit IFU. If the judgment result is no, an exception message is generated and reported. If the judgment result is yes, the next detection of the instruction decoding unit IDU is carried out; after decomposing the instructions of the instruction decoding unit IDU into multiple micro-operations, check whether the instructions before decomposition correspond to the decomposed micro-operations and whether the number of decomposed micro-operations is correct. If the judgment result is no, information is generated and reported according to the error type. If the judgment result is yes, the decomposed micro-operations are issued to the next level.

[0018] Preferably according to the present invention, the detection of the instruction load / store unit includes:

[0019] There are two 8-entry queues in the Instruction Load Store Unit (LSU), which respectively store the LOAD instructions and STORE instructions that have just entered the decoding stage. When an instruction is executed and committed, it is cleared from the Instruction Load Store Unit (LSU). In the case of cache hit, a memory access instruction stays in the Instruction Load Store Unit (LSU) for 9 cycles. In a four-issue processor, up to 36 instructions may enter during these 9 cycles. When memory access instructions are intensive, it is very likely to fill up the queues. At the same time, during the pipeline operation of the Instruction Load Store Unit (LSU), the memory access address calculation and virtual-to-physical address translation are also completed;

[0020] When detecting the Instruction Load Store Unit, the data of the Instruction Load Store Unit (LSU) is sent to entry3, and the main processor determines whether there is an exception in the execution of the LSU-level pipeline based on the data information in entry3;

[0021] Specifically, it includes: determining whether the instruction entering the Instruction Load Store Unit (LSU) is an illegal instruction. After the instruction enters the pipeline of the Instruction Load Store Unit (LSU), the main processor compares the information stored in its table entry with the prediction item. If it is determined that the instruction is not a legal instruction, the main processor clears the instruction and generates information for reporting;

[0022] When accessing a STORE instruction, the conversion between the physical address and the virtual address is generated. The main processor determines whether there is an address misalignment error based on the data information. If the determination result is yes, information is generated for reporting; if the determination result is no, the detection of the storage access part continues to check whether it is normal;

[0023] When detecting the storage access part, the information of the corresponding storage page emitted by the instruction is sent to the table entry to determine whether there is an exception in the storage page. The exceptions of the storage page can be divided into type A and type B. Type A refers to the lag in filling the storage page due to the emission queue being full; type B refers to the failure to fill the storage page or the loss of the storage page due to special errors. The main processor determines whether this exception belongs to type A or type B based on the obtained table entry data.

[0024] According to the preferred embodiment of the present invention, the detection of the Instruction Retirement Unit (RTU) includes:

[0025] After the Instruction Retirement Unit (RTU) receives the instruction retirement signal, the detection starts. The information of the Reorder Buffer (ROB) is first sent to entry4 to detect whether the micro-operation number sent to the Reorder Buffer (ROB) is correct. If the determination result is no, an exception information is generated for reporting;

[0026] The main processor determines whether the instruction about to retire in the reorder buffer ROB is retired in order according to the prediction item. If the judgment result is no, an exception message is generated and reported; if the judgment result is yes, the next item is continuously detected; the number of idle physical registers in the instruction retirement unit RTU is also sent to the table entry to determine whether the instruction retirement unit RTU can provide enough registers. If the judgment result is yes, the pipeline continues to execute; if the judgment result is no, an exception message is generated and reported, and at the same time, the main processor sends a signal to trigger the pipeline stall.

[0027] In addition, the main processor also detects the situation of the instruction in the storage cache; after the instruction is retired, it is judged whether the retired instruction is normally written back to the memory from the storage cache according to the data information in the table entry. If the judgment result is yes, there is no exception in the RTU stage pipeline of the instruction retirement unit; if the judgment result is no, an exception occurs in the write-back stage, the main processor generates an exception message and reports it, and at the same time, it may transmit a signal to other data memories to find data from other data memories.

[0028] Further preferably, each sub-table stores four pieces of information: virtual address, physical address, page size, and page attribute.

[0029] According to the preference of the present invention, an interface is provided. This interface is connected to the table entry and the processor. Peripherals are set through the interface, and the running situation of the cpu pipeline is viewed through the peripherals; when an exception occurs in the cpu pipeline, information is automatically generated and reported according to the analysis results of various data. The user can read the reason for the exception through this interface or process the data manually through this interface.

[0030] According to the preference of the present invention, the main processor collects the data corresponding to different positions of the cpu pipeline according to the data corresponding to different detection items in the hierarchical table entry, and judges whether each position of the cpu pipeline is running normally; if an abnormal situation occurs in the operation of the cpu pipeline, the exception information is reported.

[0031] Further preferably, the main processor also calculates the duty cycle of the abnormal time, analyzes the current running mode of the pipeline according to this duty cycle value. If the duty cycle reaches the threshold of 30%, a warning message is sent to the user, and the user views the pipeline information and makes an optimized design.

[0032] The beneficial effects of the present invention are as follows:

[0033] 1. The present invention provides a method for monitoring the execution of instructions at all levels of the pipeline, which can observe the execution situation of the pipeline; it can collect and analyze the data at all levels of the pipeline, so that when the pipeline runs abnormally, it can be quickly detected and sent to the user, and the pipeline exception can be found as soon as possible.

[0034] 2. The present invention provides a hierarchical table entry for storing information of different pipeline stages and classifying different pipeline stages. By creating the hierarchical table entry, data information at each stage of the pipeline is collected. Through the analysis of each item of data in the table entry, it is determined whether the pipeline stalls and whether the instruction execution is normal. If the judgment result is yes, an abnormal signal is returned to the main processor, and the main processor makes a reaction to perform data processing according to different positions of the pipeline stall or different reasons for the instruction exception. It can quickly detect when the pipeline runs abnormally and send it to the user, and as much as possible ensure the real-time detection of the pipeline exception.

[0035] 3. The pipeline instruction execution monitoring method provided by the present invention can have better interaction with users. Brief Description of the Drawings

[0036] Figure 1 is a schematic flowchart of cpu pipeline instruction execution detection in an embodiment of the present invention;

[0037] Figure 2 is a schematic diagram of the main processor determining whether the pipeline runs normally in an embodiment of the present invention. Detailed Embodiments

[0038] To better illustrate the characteristics of the present invention, the following will elaborate on the detailed embodiments in combination with specific embodiments and the drawings of the specification, but not limited thereto.

[0039] Embodiment 1

[0040] A method for monitoring cpu pipeline instruction execution, as Figure 1 shown, includes:

[0041] Obtain data at each stage of the cpu pipeline; the data at each stage of the cpu pipeline specifically includes: the number of instructions to be issued, the number of received instructions, the number of micro-operations into which the instructions are decomposed, the number of micro-operations issued, and the number of instructions or micro-operations received by the next stage; the user can manually input the desired preset detection items and store them in the main processor. These detection boxes can include detection items for each stage of the pipeline. Subsequently, the main processor can analyze this data to determine whether the cpu pipeline runs normally. This judgment method has a low misjudgment rate. At the same time, there can also be sub-detection items below each stage of the pipeline to facilitate the user to specifically view the specific module where the pipeline has an abnormality according to their own needs.

[0042] Create a hierarchical table entry to collect data at each stage of the cpu pipeline, and different detection items correspond to different positions of the cpu pipeline;

[0043] Based on the data of each level of the current CPU pipeline obtained from the hierarchical table entry, determine whether the execution of the CPU pipeline is abnormal. If the judgment result is abnormal, return the abnormal signal to the main processor, and the main processor will make a response and perform data processing according to different positions of the CPU pipeline stall or different reasons for instruction exceptions.

[0044] Embodiment 2

[0045] A method for monitoring the execution of CPU pipeline instructions according to Embodiment 1, characterized in that:

[0046] The hierarchical table entry includes multiple sub-tables, and the multiple sub-tables are divided into table entries of different levels according to the number of pipeline stages in the CPU pipeline. For example, the 1st - 4th level CPU pipeline corresponds to the first-level table entry, the 5th - 7th level CPU pipeline corresponds to the second-level table entry, the 8th - 10th level CPU pipeline corresponds to the third-level table entry, and the 11th - 14th level CPU pipeline corresponds to the fourth-level table entry. According to different positions of the CPU pipeline judged by the detection items, each level of table entry receives data of the pipeline at different stages. The main processor can then make a judgment based on the data transmitted from different table entries and analyze what kind of problem occurs in which level of the pipeline. Users can also view the running status of the CPU pipeline based on this.

[0047] Each sub-table includes multiple detection items, which respectively process data at different positions of the CPU pipeline corresponding to the number of pipeline stages; after the signals from the CPU pipeline reach the table entry, they are classified according to different positions of the source, which can reduce the time for the table entry to classify and send data to the main processor, thereby improving the performance of the CPU.

[0048] The detection items include: detection of the front end of the CPU pipeline and detection of the back end of the CPU pipeline; the detection of the front end of the pipeline includes detection of the instruction fetch unit and detection of the instruction decoding unit; the detection of the back end of the pipeline includes detection of the instruction load / store unit and detection of the instruction retirement unit RTU.

[0049] The detection of the instruction fetch unit includes:

[0050] The instruction fetch unit IFU includes the IF1 stage, the IF2 stage, the IF3 stage, and the IF4 stage. The IF1 stage is the instruction fetch stage, the IF2 stage is the instruction renaming stage, the IF3 stage is the instruction dispatch stage, and the IF4 stage is the instruction issue stage; in the IF3 stage, it is responsible for packing and preprocessing instructions, and then sending them to the instruction decoding unit IDU for decoding various instructions; at this time, such as Figure 2As shown, the flag signals corresponding to the instruction fetch unit IFU are transmitted to the entry instruction table entries, which include entry1, entry2, entry3, and entry4, and store the information collected from the pipeline; the information related to the instruction fetch unit IFU is stored in entry1; specifically, it includes information such as the fetched instruction information and the number of fetched instructions; then, the main processor compares the information in the table entry with the previously set detection items to determine whether the pipeline operation at this stage is normal. The flag signals corresponding to the IFU stage are transmitted to entry1, and the information related to the IFU is stored in entry1. Then, the main processor compares the information in the table entry with the previously set detection items to determine whether the pipeline operation at this stage is normal.

[0051] Specifically, the table entries corresponding to the instruction fetch unit IFU receive and store information including the fetched instruction information, the number of fetched instructions, etc. The main processor determines whether the fetched instructions are correct according to the original prediction items. If not, an exception occurs in the instruction fetching, and information is generated and reported instead.

[0052] The instruction fetch unit (IFU) features low power consumption, high branch prediction accuracy, and high instruction prefetch efficiency, and realizes functions such as instruction fetching, branch prediction, and jump. The IFU has a total of four pipeline stages, including three instruction fetch pipelines and one PC value calculation pipeline. The IFU designs caches such as a branch history table, a cascaded branch target buffer, and an indirect jump target buffer, and cooperates with the branch and transfer predictors to solve the problems of pipeline flushing, blocking the subsequent pipeline stages, and affecting the pipeline performance caused by the low accuracy of the branch predictor.

[0053] The detection of the instruction decoding unit includes:

[0054] The instruction decoding unit IDU decodes four instructions simultaneously. After completing the detection of data dependencies, decomposition of micro-operations, register mapping, instruction dispatching, and instruction issuing processes, the instructions are issued to the next-level CPU pipeline. At this time, the data corresponding to the instruction decoding unit IDU is transmitted to entry2, and the main processor compares the data in entry2 with the detection items to determine whether the pipeline instructions of the IDU stage are executed normally;

[0055] Specifically, it includes: checking whether the instructions entering the Instruction Decode Unit (IDU) correspond one-to-one with the instructions fetched and sent by the Instruction Fetch Unit (IFU). If the judgment result is negative, an exception message is generated and reported. If the judgment result is positive, the next-step detection of the IDU is carried out; after decomposing the instructions of the IDU into multiple micro-operations, checking whether the instructions before decomposition correspond to the decomposed micro-operations and whether the number of decomposed micro-operations is correct. If the judgment result is negative, information is generated and reported according to the error type. If the judgment result is positive, the decomposed micro-operations are dispatched to the next level.

[0056] The instruction items to be tested can be uploaded to the main processor in advance by the user. The number of micro-operations decomposed from different instructions will be different, but the detection methods are similar. After the detection is completed, if an exception occurs, the information can be sent to the user through the interface so that the user can make a solution.

[0057] The Instruction Decode Unit (IDU) decodes four instructions simultaneously, detects data dependencies, decomposes micro-operations, performs register mapping, instruction dispatching, and instruction emission. The 4 instructions fetched per cycle are temporarily stored in a 128-entry instruction cache queue, which is implemented by a 4-port FIFO queue. Four decode units decode the four instructions simultaneously, cache the decoded micro-operations in the micro-operation queue waiting for subsequent pipelining and scheduling, and the instruction IDs need to be written into the Reorder Buffer (ROB). At the same time, 32 logical registers are mapped to 128 physical registers, the register mapping table is updated, and data hazards are eliminated.

[0058] The detection of the Instruction Load / Store Unit includes:

[0059] There are two 8-entry queues in the Instruction Load / Store Unit (LSU), which respectively store the LOAD instructions and STORE instructions that have just entered the decode stage. When an instruction is executed and committed, the instruction is cleared from the LSU. Therefore, in the case of cache hits, a memory access instruction stays in the LSU for 9 cycles. In a four-issue processor, at most 36 instructions may enter in 9 cycles, and it is very likely to fill up the queue when memory access instructions are intensive. At the same time, the memory access address calculation and virtual-to-physical address translation are also completed during the pipeline operation of the LSU;

[0060] When detecting the Instruction Load / Store Unit, the data of the LSU is sent to entry3, and the main processor judges whether the LSU-level pipeline execution is abnormal through the data information in entry3;

[0061] Specifically, it includes: determining whether the instruction entering the instruction load / store unit (LSU) is an illegal instruction. After this instruction enters the LSU pipeline, the main processor compares the information stored in the table entry with the prediction entry. If it is determined that the instruction is not a legal instruction, the main processor clears the instruction and generates information for reporting.

[0062] When accessing a STORE instruction, the conversion between the physical address and the virtual address is generated, which is completed by the memory management unit. However, the converted virtual address may not be in the same page, and in this case, an address misalignment error may occur. This address information will be stored in the table entry. The main processor determines whether an address misalignment error has occurred based on the data information. If the determination result is yes, information is generated for reporting; if the determination result is no, the main processor continues to detect whether the storage access part is normal.

[0063] When detecting the storage access part, the information of the corresponding storage page emitted by the instruction is sent to the table entry to determine whether the storage page has an exception. It should be noted that the exceptions of the storage page can be divided into type A and type B. Type A refers to the storage page filling lag caused by the emission queue being full; type B refers to the storage page not being filled or the storage page being lost due to special errors. The main processor determines whether this exception belongs to type A or type B based on the obtained table entry data.

[0064] The instruction load / store unit (LSU) supports double-issue of scalar store / load instructions, single-issue of vector store / load instructions, and full out-of-order execution of all store / load instructions, and supports non-blocking access to the cache. It supports store / load instructions of bytes, half-words, words, double-words, and quad-words, and supports sign extension and zero extension of load instructions of bytes and half-words. The store / load instructions can be pipelined, enabling a data throughput of accessing one data per cycle. The design supports 8-way data stream prefetching, and the data is put into the L1 data cache in advance.

[0065] The detection of the instruction retirement unit (RTU) includes:

[0066] When an instruction has an exception, this instruction must be retired. The instruction retirement unit (RTU) is responsible for the write-back and retirement of instructions, including a reorder buffer (ROB) and a physical register file. Among them, the reorder buffer (ROB) is responsible for out-of-order recovery and in-order retirement of instructions, and the physical register file is responsible for out-of-order recovery and transfer of results. The instruction retirement unit retires four instructions in parallel per clock cycle and releases physical registers.

[0067] After the instruction retirement unit (RTU) receives the instruction retirement signal, the detection starts; the information of the reorder buffer (ROB) is first sent to entry4, and it is detected whether the micro-operation numbers sent to the reorder buffer (ROB) are correct. If the determination result is no, exception information is generated for reporting.

[0068] The main processor determines whether the instruction about to retire in the reorder buffer (ROB) is to be retired in order according to the prediction item. If the judgment result is no, an exception message is generated and reported; if the judgment result is yes, the next item is continuously detected; the number of free physical registers in the instruction retirement unit (RTU) is also sent to the table entry to determine whether the instruction retirement unit (RTU) can provide enough registers. If the judgment result is yes, the pipeline continues to execute; if the judgment result is no, an exception message is generated and reported, and at the same time, the main processor sends a signal to trigger a pipeline stall.

[0069] In addition, the main processor also detects the situation of the instruction in the store buffer; after the instruction is retired, it is judged whether the retired instruction is normally written back to the memory from the store buffer according to the data information in the table entry. If the judgment result is yes, there is no exception in the RTU-level pipeline of the instruction retirement unit; if the judgment result is no, an exception occurs in the write-back stage, the main processor generates an exception message and reports it, and at the same time, it may transmit a signal to other data memories to find data from other data memories.

[0070] The instruction retirement unit (RTU) is responsible for the write-back and retirement of instructions, including a reorder buffer (ROB) and a physical register file. Among them, the reorder buffer is responsible for the out-of-order recovery and in-order retirement of instructions, and the physical register file is responsible for the out-of-order recovery and transfer of results. By supporting the parallel recovery and fast retirement of instructions, the instruction retirement efficiency is improved. The instruction retirement unit retires four instructions in parallel every clock cycle and releases physical registers.

[0071] Each sub-table stores four pieces of information: virtual address, physical address, page size, and page attribute.

[0072] It is judged whether the result extracted by the instruction meets the exception determination conditions detected at the front end of the pipeline; if the determination result is no, it is judged whether the decoded instruction corresponds to the decomposed micro-operations; in addition, it is necessary to judge whether the data received by the load / store unit corresponds to the emitted micro-operations. In this way, a specific implementation form for detecting the execution of the cpu pipeline is provided.

[0073] Embodiment 3

[0074] According to the method for monitoring the execution of cpu pipeline instructions described in Embodiment 2, the difference is that:

[0075] An interface is provided, which is connected to the table entry and the processor. Peripherals are set through the interface, and users can view the running status of the cpu pipeline by connecting reasonable peripherals; when an exception occurs in the cpu pipeline, information is automatically generated and reported according to the analysis results of various data, and users can read the reason for the exception through this interface or process the data manually through this interface.

[0076] The main processor collects the data at different positions of the corresponding CPU pipeline according to the data at different positions of the CPU pipeline corresponding to different detection items in the hierarchical table entry, and determines whether each position of the CPU pipeline is operating normally; if an abnormal situation occurs in the operation of the CPU pipeline, the abnormal information is reported.

[0077] For the flag signal corresponding to the IFU level, it will be transmitted to entry1, and entry1 will store the information related to the IFU. Then the main processor can compare the information in the table entry with the previously set detection items to determine whether the pipeline operation at this stage is normal.

[0078] Idu: First, it checks whether the instructions entering the IDU level correspond one by one to the instructions extracted and sent at the IFU level. If the judgment result is no, an abnormal information is generated and reported. If the judgment result is yes, the next step of detection at the IDU level is carried out; after decomposing the instructions into multiple micro-operations, it checks whether the instructions before decomposition correspond to the decomposed micro-operations and whether the number of decomposed micro-operations is correct. If the judgment result is no, information is generated and reported according to the error type. If the judgment result is yes, the decomposed micro-operations are issued to the next level.

[0079] LSU: First, it judges whether the instructions entering the LSU level are illegal instructions. After the instructions enter the LSU-level pipeline, the main processor can compare the information stored in the table entry with the prediction item. If it is judged that the instruction is not a legal instruction, the main processor will clear the instruction and generate information for reporting.

[0080] RTU: The main processor judges whether the instructions about to retire in the ROB are retired in sequence according to the prediction item. If the judgment result is no, an abnormal information is generated and reported; if the judgment result is yes, the next item is continuously detected; the number of idle physical registers in the RTU is also sent to the table entry to judge whether the RTU can provide enough registers. If the judgment result is yes, the pipeline continues to execute. If the judgment result is no, an abnormal information is generated and reported, and at the same time the main processor issues a signal to trigger a pipeline stall.

[0081] In addition, the situation of instructions in the store cache is also detected. After the instruction retires, it is judged whether the retired instruction is normally written back to the memory from the store cache through the data information in the table entry. If the judgment result is yes, there is no abnormality in the RTU-level pipeline. If the judgment result is no, an abnormality occurs in the write-back stage, and the main processor generates an exception message for reporting. At the same time, signals may be transmitted to other data memories to search for data in other data memories. The main processor can collect and analyze the information transmitted from the pipeline as soon as possible to monitor the operation of the CPU pipeline, ensure the timeliness of the feedback of the exception information, and be closer to the actual situation.

[0082] The main processor can monitor multiple table entries and multiple-level CPU pipelines simultaneously. The main processor also calculates the duty cycle of the abnormal time. According to this duty cycle value, it analyzes the current operation mode of the pipeline. If the duty cycle reaches the threshold of 30%, when the pipeline gets stuck, it will occupy clock cycles. The number of blocked clock cycles occupied within a certain period of time divided by the total number of clock cycles during this period is called the duty cycle of the abnormal time. Then a warning message is sent to the user, and the user can view the pipeline information and make an optimized design.

Claims

1. A method for monitoring the execution of CPU pipeline instructions, characterized in that Including: Obtain data at all levels of the CPU pipeline; Create hierarchical table entries to collect data at all levels of the CPU pipeline. Different detection items correspond to different positions in the CPU pipeline; Based on the current data at all levels of the CPU pipeline obtained from the hierarchical table entries, determine whether the execution of the CPU pipeline is abnormal. If the judgment result is abnormal, return the abnormal signal to the main processor. The main processor makes a reaction and processes the data according to different positions of the CPU pipeline stall or different reasons for instruction exceptions; The hierarchical table entries include multiple sub-tables, and the multiple sub-tables are divided into table entries of different levels according to the number of pipeline stages in the CPU pipeline; Each sub-table includes multiple detection items, and respectively processes data at different positions of the CPU pipeline corresponding to the number of pipeline stages; The detection items include: detection at the front end of the CPU pipeline and detection at the back end of the CPU pipeline; Detection at the front end of the pipeline includes detection of the instruction fetch unit and detection of the instruction decoding unit; Detection at the back end of the pipeline includes detection of the instruction load / store unit and detection of the instruction retirement unit RTU.

2. The method for monitoring the execution of CPU pipeline instructions according to claim 1, wherein The detection of the instruction fetch unit includes: The instruction fetch unit IFU includes the IF1 stage, the IF2 stage, the IF3 stage, and the IF4 stage. The IF1 stage is the instruction fetch stage, the IF2 stage is the instruction renaming stage, the IF3 stage is the instruction dispatch stage, and the IF4 stage is the instruction issue stage; In the IF3 stage, it is responsible for packing and preprocessing instructions, and then sending them to the instruction decoding unit IDU for decoding various instructions; At this time, the flag signal corresponding to the instruction fetch unit IFU is transmitted to the entry finger table entry. The entry finger table entry includes entry1, entry2, entry3, and entry4, which store information collected from the pipeline; Information related to the instruction fetch unit IFU is stored in entry1; After that, the main processor compares the information in the table entry with the previously set detection items to determine whether the pipeline operation at this stage is normal.

3. A method for monitoring the execution of CPU pipeline instructions according to claim 1, wherein The detection of the instruction decoding unit includes: The instruction decoding unit IDU decodes four instructions simultaneously. After completing the detection of data dependencies, decomposition of micro-operations, register mapping, instruction dispatch, and instruction issue processes, it issues the instructions to the next level of the CPU pipeline. At this time, the data corresponding to the instruction decoding unit IDU is transmitted to entry2, and the main processor compares the data in entry2 with the detection items to determine whether the instruction execution of the IDU-level pipeline is normal; Specifically including: checking whether the instructions entering the instruction decoding unit IDU correspond one by one to the instructions fetched and sent by the instruction fetch unit IFU. If the judgment result is no, generate an abnormal message for reporting. If the judgment result is yes, perform the next detection of the instruction decoding unit IDU; After decomposing the instructions of the instruction decoding unit IDU into multiple micro-operations, check whether the instructions before decomposition correspond to the decomposed micro-operations and whether the number of decomposed micro-operations is correct. If the judgment result is no, generate information for reporting according to the error type. If the judgment result is yes, emit the decomposed micro-operations to the next level.

4. A method for monitoring the execution of CPU pipeline instructions according to claim 1, characterized in that, Detection of the instruction load / store unit, including: There are two 8-entry queues in the instruction load / store unit LSU, which respectively store the LOAD instructions and STORE instructions that have just entered the decoding stage. When an instruction is executed and committed, it is cleared from the instruction load / store unit LSU; in the case of cache hit, a memory access instruction stays in the instruction load / store unit LSU for 9 cycles; meanwhile, during the pipeline operation of the instruction load / store unit LSU, memory access address calculation and virtual-to-physical address translation are also completed; When detecting the instruction load / store unit, the data of the instruction load / store unit LSU is sent to entry3, and the main processor determines whether there is an abnormality in the execution of the LSU-level pipeline based on the data information in entry3; Specifically, it includes: determining whether the instruction entering the instruction load / store unit LSU is an illegal instruction. After the instruction enters the instruction load / store unit LSU pipeline, the main processor compares the information stored in its table entry with the prediction item. If it is determined that the instruction is not a legal instruction, the main processor clears the instruction and generates information for reporting; When accessing the STORE instruction, the conversion between the physical address and the virtual address is generated. The main processor determines whether there is an unaligned address error based on the data information. If the determination result is yes, information is generated for reporting; if the determination result is no, the detection of the storage access part continues to check whether it is normal; When detecting the storage access part, the information of the corresponding storage page emitted by the instruction is sent to the table entry to determine whether there is an abnormality in the storage page; the abnormalities of the storage page can be divided into type A and type B. Type A refers to the storage page filling lag caused by the emission queue being full; type B refers to the storage page not being filled or the storage page being lost due to special errors; the main processor determines whether this abnormality belongs to type A or type B based on the obtained table entry data.

5. A method for monitoring the execution of CPU pipeline instructions according to claim 1, characterized in that Detection of the instruction retirement unit RTU, including: After the instruction retirement unit RTU receives the instruction retirement signal, the detection starts; the information of the reorder buffer ROB is first sent to entry4, and it is detected whether the micro-operation number sent to the reorder buffer ROB is correct. If the determination result is no, abnormal information is generated for reporting; The main processor determines whether the instruction about to retire in the reorder buffer ROB is retired in sequence according to the prediction item. If the determination result is no, abnormal information is generated for reporting; if the determination result is yes, the detection of the next item continues; the number of free physical registers in the instruction retirement unit RTU is also sent to the table entry to determine whether the instruction retirement unit RTU can provide enough registers. If the determination result is yes, the pipeline continues to execute; if the determination result is no, abnormal information is generated for reporting, and at the same time, the main processor issues a signal to trigger the pipeline stall; In addition, the main processor also detects the situation of instructions in the store cache; after an instruction retires, it determines whether the retired instruction is normally written back to the memory from the store cache based on the data information in the table entry. If the judgment result is yes, there is no abnormality in the RTU-level pipeline of the instruction retirement unit. If the judgment result is no, an abnormality occurs in the write-back stage. The main processor generates an exception message and reports it. At the same time, it may transmit a signal to other data memories to search for data in other data memories.

6. The method for monitoring the execution of CPU pipeline instructions according to claim 5, wherein Each sub-table stores four pieces of information: virtual address, physical address, page size, and page attributes.

7. A method for monitoring the execution of CPU pipeline instructions according to claim 1, characterized in that, There is an interface that is connected to the table entry and the processor. The peripherals are set through the interface, and the operation of the CPU pipeline is viewed through the peripherals. When an exception occurs in the CPU pipeline, information is automatically generated and reported based on the analysis results of various data. The user can read the cause of the exception through this interface or process the data manually through this interface.

8. A method for monitoring the execution of CPU pipeline instructions according to claim 1, characterized in that, The main processor collects the data at different positions of the corresponding CPU pipeline by itself according to the data at different positions of the CPU pipeline corresponding to different detection items in the hierarchical table entry, and determines whether each position of the CPU pipeline is operating normally. If an abnormal situation occurs in the operation of the CPU pipeline, an exception message is reported.

9. A method for monitoring the execution of CPU pipeline instructions according to any one of claims 1-8, characterized in that, The main processor also calculates the duty cycle of the abnormal time. Based on this duty cycle value, it analyzes the current operation mode of the pipeline. If the duty cycle reaches the threshold of 30%, a warning message is sent to the user, and the user views the pipeline information and makes an optimized design.

Citation Information

Patent Citations

  • Pipeline-level operation device, data processing method and network-on-chip chip

    CN105468335A

  • Avoidance method for conflict between instruction sets in RISC-CPU and avoidance system thereof

    CN106610816A