Processor performance hardware optimization method and system, electronic equipment and storage medium

By monitoring the real-time performance data of all target hardware event controllers on the processor, determining performance bottlenecks and optimizing, the problems of inefficiency and low accuracy of traditional methods are solved, and more efficient performance monitoring and optimization are achieved.

CN120196527AActive Publication Date: 2025-06-24SHANDONG BOSUAN ZHIXIN INFORMATION TECHNOLOGY CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510676947.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Traditional CPU performance monitoring and optimization methods are inefficient and the monitoring results are not accurate, resulting in the inability to accurately identify performance bottlenecks and effectively optimize them.

Method used

By monitoring the real-time performance data of all target hardware event controllers on the processor, determine the degree of deviation between the real-time performance data and standard performance data, bind the hardware events that meet the threshold to the hardware event counter, determine the performance bottleneck based on the target number, and perform targeted performance optimization.

Benefits of technology

Improve the efficiency and accuracy of performance bottleneck monitoring, and more accurately identify the performance bottleneck of the processor and optimize it, thereby improving the overall performance of the processor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196527A_ABST
    Figure CN120196527A_ABST
Patent Text Reader

Abstract

The invention provides a processor performance hardware optimization method and system, electronic equipment and a storage medium, and relates to the technical field of computers. The processor performance hardware optimization method comprises the steps of controlling all target hardware event controllers on a processor to monitor real-time performance data of corresponding hardware events in response to starting of a performance monitoring function for the processor; determining a first deviation degree between the real-time performance data and the corresponding standard performance data; binding the hardware event corresponding to the first deviation degree meeting the first threshold value to a hardware event counter, and monitoring a target number of times of the corresponding hardware event based on the hardware event counter; determining a performance bottleneck of the processor based on the target number of times; and performing performance optimization on the processor based on the performance bottleneck. The performance bottleneck monitoring efficiency and precision can be improved, and the performance of the processor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, a system, an electronic device, and a storage medium for optimizing the performance of a processor in terms of hardware. Background Art

[0002] The performance monitoring of a central processing unit (CPU) requires collecting application execution information on a system. For example, hardware events are directly collected from the CPU and / or the system through hardware monitoring technologies to identify performance bottlenecks, and then the CPU is optimized in terms of kernel microarchitecture, cache architecture, program architecture, etc. based on the performance bottlenecks. To collect hardware events, modern CPUs implement dedicated performance monitoring units (PMUs) for monitoring various hardware events. Among them, a hardware event refers to a behavioral event of a hardware function inside the CPU.

[0003] The traditional methods for CPU performance monitoring and optimization are as follows: Configure the set hardware event counters (EVENT_COUNTER) to the hardware events to be monitored, start each EVENT_COUNTER and the corresponding hardware event, collect the values of the EVENT_COUNTER for numerical analysis, then monitor different hardware events multiple times, repeat the above process, draw a conclusion on the performance bottleneck, and perform CPU performance optimization based on the performance bottleneck. However, the traditional methods for CPU performance monitoring and optimization have the following defects: 1. Slow monitoring efficiency: There are up to hundreds of hardware events supported inside the CPU, but the number of EVENT_COUNTERs is small, for example, 5. Therefore, using EVENT_COUNTERs to cyclically monitor all hardware events is extremely inefficient; 2. Low accuracy of monitoring results: For example, if the traditional solution monitors hardware event A and determines that there are some bottleneck limitations in hardware event A, it concludes that the performance bottleneck of the entire CPU system lies in hardware event A. However, in fact, the real performance bottleneck lies in hardware event B that has not been monitored, which will result in inaccurate monitoring results of the performance bottleneck. Even if CPU performance optimization is performed based on the obtained monitoring results of the performance bottleneck, the best performance of the CPU cannot be truly exerted. Summary of the Invention

[0004] The present disclosure provides a method, a system, an electronic device, and a storage medium for optimizing the performance of a processor in terms of hardware, so as to at least solve the above technical problems existing in the prior art.

[0005] According to a first aspect of the present disclosure, there is provided a method for hardware optimization of processor performance, including: in response to the activation of the performance monitoring function for the processor, controlling all target hardware event controllers on the processor to monitor the real-time performance data of corresponding hardware events; determining a first deviation degree between the real-time performance data and the corresponding standard performance data; binding the hardware events corresponding to the first deviation degree that meets a first threshold to a hardware event counter, and monitoring the target number of times of the corresponding hardware events based on the hardware event counter; determining the performance bottleneck of the processor based on the target number of times; and performing performance optimization on the processor based on the performance bottleneck.

[0006] In another implementable manner, before determining the first deviation degree between the real-time performance data and the corresponding standard performance data, the method further includes: running the processor based on a standard test model program to obtain the original performance standards corresponding to all hardware events on the processor; determining a second deviation degree between the original performance standards and the standard performance data; in response to the second deviation degree meeting a second threshold, generating warning data of abnormal standard performance data and outputting it to the client; and obtaining new standard performance data; the new standard performance data is the standard performance data regenerated by the client based on the warning data, or the standard performance data obtained by weighted averaging the original performance standards and the standard performance data.

[0007] In another implementable manner, determining the first deviation degree between the real-time performance data and the corresponding standard performance data includes: determining the difference between the real-time performance data and the corresponding standard performance data; and determining the ratio of the difference to the standard performance data as the first deviation degree.

[0008] In another implementable manner, performing performance optimization on the processor based on the performance bottleneck includes: in response to the performance bottleneck existing in the core microarchitecture of the processor, monitoring the processing duration of each processing stage in the instruction processing flow of the core microarchitecture; and performing corresponding optimization measures on the processing stage whose processing duration meets a third threshold.

[0009] In another implementable manner, monitoring the processing duration of each processing stage in the instruction processing flow of the core microarchitecture includes: monitoring the interval duration between two instruction fetches in the instruction fetch stage; monitoring the interval duration between two decodings in the decoding stage; monitoring the duration for each execution unit to complete instruction data operation in the execution stage; monitoring the duration of the memory access operation for each instruction in the memory access stage; and monitoring the duration for each instruction to be written back to the target storage location of the processor in the write-back stage.

[0010] In another feasible implementation, corresponding optimization measures are executed for the processing stage whose processing duration meets the third threshold, including: in response to the processing durations of the instruction fetch stage and the instruction decode stage meeting the third threshold, enabling the instruction prefetch function, increasing the instruction issue width, and / or optimizing the mapping relationship of the instruction translation lookaside buffer (ITLB); the ITLB stores the mapping relationship between the virtual address and the physical address corresponding to the instruction; in response to the processing durations of the execution stage and the memory access stage meeting the third threshold, increasing the instruction execution width and / or optimizing the mapping relationship of the data translation lookaside buffer (DTLB); the DTLB stores the mapping relationship between the virtual address and the physical address corresponding to the data; in response to the processing duration of the write-back stage meeting the third threshold, increasing the cache depth of the reorder buffer (ROB) and / or performing fast processing of the write-back status; the fast processing of the write-back status includes early write-back, enhanced register renaming, and / or increasing the write-back bandwidth.

[0011] In another feasible implementation, the performance optimization of the processor based on the performance bottleneck includes: in response to the performance bottleneck existing in the cache architecture of the processor, monitoring the data hit rates corresponding to different cache parts in the cache architecture; the cache parts include the translation lookaside buffer (TLB), the level-1 cache, and the level-2 cache; and executing corresponding optimization measures for the cache part whose data hit rate meets the fourth threshold.

[0012] In another feasible implementation, the execution of corresponding optimization measures for the cache part whose data hit rate meets the fourth threshold includes: in response to the data hit rate of the TLB meeting the fourth threshold, updating the page table storing the mapping relationship between the virtual address and the physical address in the double data rate (DDR) memory to the TLB based on the rules for the software to send virtual addresses; in response to the data hit rate of the level-1 cache meeting the fourth threshold, updating the corresponding access address in the level-2 cache or the DDR memory to the level-1 cache based on the rules for the system bus to send addresses; in response to the data hit rate of the level-2 cache meeting the fourth threshold, updating the corresponding access address in the DDR memory to the level-2 cache based on the rules for the system bus to send addresses.

[0013] According to a second aspect of the present disclosure, there is provided a processor performance hardware optimization system, including: a bottleneck analysis controller, configured to, in response to enabling a performance monitoring function for a processor, control all target hardware event controllers on the processor to monitor real-time performance data of corresponding hardware events; determine a first deviation degree between the real-time performance data and corresponding standard performance data; bind the hardware events corresponding to the first deviation degree that meets a first threshold to a hardware event counter, and monitor a target number of corresponding hardware events based on the hardware event counter; determine a performance bottleneck of the processor based on the target number; and a performance optimizer, configured to perform performance optimization on the processor based on the performance bottleneck.

[0014] In another implementable manner, the bottleneck analysis controller is further configured to: run the processor based on a standard test model program to obtain original performance standards corresponding to all hardware events on the processor; determine a second deviation degree between the original performance standards and the standard performance data; in response to the second deviation degree meeting a second threshold, generate warning data of abnormal standard performance data and output it to a user terminal; and obtain new standard performance data; the new standard performance data is standard performance data regenerated by the user terminal based on the warning data, or standard performance data obtained by performing weighted averaging on the original performance standards and the standard performance data.

[0015] In another implementable manner, the bottleneck analysis controller is further configured to: determine a difference between the real-time performance data and corresponding standard performance data; and determine a ratio of the difference to the standard performance data as the first deviation degree.

[0016] In another implementable manner, the performance optimizer includes a microarchitecture controller, and the microarchitecture controller is configured to: in response to the performance bottleneck existing in the core microarchitecture of the processor, monitor a processing duration of each processing stage in an instruction processing flow of the core microarchitecture; and perform corresponding optimization measures on the processing stage whose processing duration meets a third threshold.

[0017] In another implementable manner, the microarchitecture controller is further configured to: monitor an interval duration between two instruction fetches in an instruction fetch stage; monitor an interval duration between two decodings in a decoding stage; monitor a duration for each execution unit to complete instruction data operation in an execution stage; monitor a duration of a memory access operation for each instruction in a memory access stage; and monitor a duration for each instruction to be written back to a target storage location of the processor in a write-back stage.

[0018] According to a third aspect of the present disclosure, there is provided an electronic device, including: the processor performance hardware optimization system described in the present disclosure.

[0019] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the present disclosure.

[0020] In the method, system, electronic device, and storage medium for optimizing the performance of a processor according to the present disclosure, when monitoring the performance of the processor, real-time performance data of all hardware events is synchronously monitored to obtain all-round performance bottlenecks of the processor, which can improve the efficiency and accuracy of performance bottleneck monitoring; moreover, based on the performance bottlenecks, targeted optimizations are made to the core microarchitecture and cache architecture of the processor, which can improve the performance of the processor.

[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become easily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, where: In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.

[0023] Figure 1 shows a flowchart of a method for optimizing the performance of a processor according to an embodiment of the present disclosure Figure 1 ; Figure 2 shows a flowchart of a method for optimizing the performance of a processor according to an embodiment of the present disclosure Figure 2 ; Figure 3 shows a flowchart of a method for optimizing the performance of a processor according to an embodiment of the present disclosure Figure 3 ; Figure 4 shows a schematic structural diagram of a processor performance optimization system in the prior art; Figure 5 shows a schematic structural diagram of a system for optimizing the performance of a processor according to an embodiment of the present disclosure; Figure 6 shows a schematic structural diagram of a cache architecture controller according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] To make the objectives, features, and advantages of the present disclosure more apparent and understandable, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present disclosure.

[0025] Figure 4 shows a schematic structural diagram of a prior-art processor performance optimization system, as Figure 4 shown, the traditional CPU main structure includes a processing core Core. If it is a multi-core processor, there will be multiple cores, that is, Figure 4 Core0 - CoreN in

[0026] Core0 - CoreN in the figure. The processor core is connected to the first-level cache (L1 Cache) and the second-level cache (L2 Cache). The L1 Cache and L2 Cache form the cache system of the CPU. The L2 Cache is connected to the double data rate (DDR, Double Data Rate) memory. The performance monitoring controller (PMC, Performance Monitoring Controller) is used to monitor the performance of the entire system. The hardware event controller (Event_Ctrl) has a large amount of data and is distributed in each functional module of the CPU, and is used to monitor the hardware events on different hierarchical structures of the CPU. For example, the hardware event controller at the processor core level can monitor the instructions per cycle (IPC, Instructions Per Cycle) within the core, the pipeline front-end stall efficiency, the pipeline back-end stall efficiency, etc. The hardware event controller at the L1 Cache level can monitor the L1 Cache hit rate, load / store (Load / Store) memory operand, etc.

[0027] Figure 1 shows a flowchart of a processor performance hardware optimization method according to an embodiment of the present disclosure, as Figure 1 shown, a processor performance hardware optimization method includes: Figure 1 As shown in the figure, a processor performance hardware optimization method includes: Step S101, in response to the activation of the performance monitoring function for the processor, control all target hardware event controllers on the processor to monitor the real-time performance data of the corresponding hardware events.

[0028] In this embodiment, when the user or the system activates the performance monitoring function of the processor, all target hardware event controllers on the processor are activated. The target hardware event controller is responsible for monitoring the real-time performance data of the corresponding hardware event. For example, hardware event controller A monitors the hit rate of the L1 Cache, and hardware event controller B monitors the instructions per cycle (IPC), etc. Among them, the processor can be a RISC-V CPU, and the RISC-V CPU is a CPU based on multi-core reduced instruction set computing (RISC). The target hardware event controller can be some or all of the hardware event controllers in the processor. The hardware event controllers corresponding to the hardware events that the user cares about can be determined as the target hardware event controllers, and the hardware event controllers corresponding to the hardware events that the user does not want to care about can be masked and the results of the masked hardware event controllers will not be received.

[0029] In the traditional solution, the hardware event controller is implemented to monitor the corresponding hardware event. Only after the software issues the configuration and enables the corresponding hardware event controller, the hardware event controller starts to work. However, in this embodiment, after the software enables the PMC of the processor, all hardware event controllers on the processor start to work, and the monitoring time (T_Monitor) of the hardware event controller can be agreed upon. The monitoring time can be flexibly configured by the user. After the monitoring duration reaches T_Monitor, all hardware event controllers stop working.

[0030] Step S102, determine the first deviation degree between the real-time performance data and the corresponding standard performance data.

[0031] In this embodiment, the real-time performance data (Real_performance) refers to the actual performance of the current hardware event collected by the hardware event controller during the monitoring process. The standard performance data (Basic_performance) is the performance benchmark that is preset by the user or the system and expected to be achieved by the hardware event. The difference between the real-time performance data and the standard performance data can be calculated, and this difference is determined as the first deviation degree. For example, if the standard performance data of hardware event A is 90% and the real-time performance data is 85%, then the first deviation degree is 90% - 85%. The first deviation degree reflects the gap between the actual performance and the expected performance of the current hardware event, providing a quantitative basis for subsequent performance bottleneck judgment. The standard performance data corresponding to different hardware events can be different. For example, if the hardware event controller C monitors the L1 hit rate and the hardware event controller D monitors the IPC, the standard performance data of the hardware event controller C can be configured as 90%, while the standard performance data of the hardware event controller D can be configured as 3.

[0032] Step S103: Bind the hardware event corresponding to the first deviation degree that meets the first threshold to the hardware event counter, and monitor the target number of times of the corresponding hardware event based on the hardware event counter.

[0033] In this embodiment, the first threshold is a preset value used to screen out those hardware events with a large gap between the actual performance and the standard performance. When the first deviation degree of a certain hardware event exceeds the first threshold, it indicates that there is a significant gap between the real-time performance of this hardware event and the expectation, and it may be a potential source of performance bottleneck. Then, this hardware event can be bound to the hardware event counter (EVENT_COUNTER). The hardware event counter is used to record the target number of times of this hardware event within a certain period of time, that is, the frequency of occurrence of this hardware event. The target number of times represents the number of times that the real-time performance data of the corresponding hardware event does not meet the standard performance data, and can be used to analyze the performance bottleneck.

[0034] Step S104: Determine the performance bottleneck of the processor based on the target number of times.

[0035] In this embodiment, according to the target number of times recorded by the hardware event counter, it can be determined which hardware events occur frequently and have poor performance. The hardware event with a high target number of times is more likely to be the location of the processor performance bottleneck. For example, if the target number of times of hardware event A is much higher than that of other hardware events and its first deviation degree is also large, it can be determined that hardware event A is the performance bottleneck of the current processor.

[0036] Step S105: Optimize the performance of the processor based on the performance bottleneck.

[0037] In this embodiment, after determining the performance bottleneck of the processor, corresponding optimization measures are taken according to the specific situation of the performance bottleneck. For example, if the performance bottleneck is caused by a low hit rate of the L1 Cache, the cache management strategy can be optimized, such as adjusting the cache size, improving the cache replacement algorithm, etc.; if the performance bottleneck is caused by low instruction execution efficiency, the instruction scheduling can be optimized, the depth of the instruction pipeline can be increased, etc. By optimizing the performance bottleneck, the overall performance of the processor can be effectively improved, making it better meet the application requirements.

[0038] In the present disclosure, hardware self-monitoring of the processor performance bottleneck is implemented, and the sorting of the performance bottlenecks of all hardware events is obtained, thus finding the real performance bottleneck on the processor. Compared with the traditional solution, it can greatly accelerate the speed of performance monitoring and improve the accuracy of performance monitoring.

[0039] In another embodiment, before step S102, "determine the first deviation degree between the real-time performance data and the corresponding standard performance data", a hardware optimization method for processor performance further includes: Run the processor based on the standard test model program to obtain the original performance standards corresponding to all hardware events on the processor.

[0040] In this embodiment, in order to ensure the accuracy of the standard performance data, the present invention uses a standard test model program to run the processor. The standard test model program is a representative test program that can comprehensively cover various functions of the processor. By running this program, the original performance standards (Basic_performance_model) of all hardware events on the processor under standard test conditions can be obtained. For example, after running the standard test model program, the original performance standard of hardware event A is 88%, and the original performance standard of hardware event B is 3.2, etc.

[0041] Determine the second deviation degree between the original performance standard and the standard performance data.

[0042] In this embodiment, the original performance standard obtained by the standard test model program is compared with the standard performance data preset by the user or the system, the difference between them is calculated, and this difference is determined as the second deviation degree. For example, if the original performance standard of hardware event A is 88% and the standard performance data set by the user is 90%, then the second deviation degree is 90% - 88%. The second deviation degree reflects the difference between the standard performance data set by the user and the original performance standard obtained from the actual test, and is used to judge the rationality of the standard performance data.

[0043] In response to the second deviation degree meeting the second threshold, generate warning data of abnormal standard performance data and output it to the user terminal.

[0044] In this embodiment, the second threshold is a preset value used to determine whether the standard performance data is abnormal. When the second deviation degree exceeds the second threshold, it indicates that there is a large difference between the standard performance data set by the user and the original performance standard obtained from the actual test, which may be caused by unreasonable user settings or changes in the test environment. At this time, warning data indicating abnormal standard performance data is generated and output to the user terminal to remind the user that there may be a problem with the standard performance data and that it needs to be re-evaluated or adjusted.

[0045] Obtain new standard performance data; the new standard performance data is the standard performance data regenerated by the user terminal based on the warning data, or the standard performance data obtained by weighted averaging the original performance standard and the standard performance data.

[0046] In this embodiment, after receiving the warning data, the user terminal can regenerate the standard performance data according to the actual situation. The new standard performance data can be set by the user again according to new test results or experience, or obtained by weighted averaging the original performance standard and the standard performance data set by the user. For example, if the user believes that the original performance standard is closer to the actual operating conditions, the original performance standard can be used as the new standard performance data, or a new standard performance data can be issued again; if the user has not issued new standard performance data for a long time, the original performance standard and the standard performance data set by the user can be weighted averaged to obtain a new standard performance data. In this way, the accuracy and rationality of the standard performance data can be ensured, providing a reliable basis for subsequent performance monitoring and optimization.

[0047] In another embodiment, step S102, "determine the first deviation degree between the real-time performance data and the corresponding standard performance data", includes: Determine the difference between the real-time performance data and the corresponding standard performance data; determine the ratio of the difference to the standard performance data as the first deviation degree.

[0048] In this embodiment, when calculating the first deviation degree, first calculate the difference between the real-time performance data and the standard performance data. For example, for hardware event A, the real-time performance data is 85% and the standard performance data is 90%, then the difference is 90% - 85% = 5%. The difference reflects the absolute gap between the real-time performance data and the standard performance data, which is the basic data for calculating the first deviation degree. Then, determine the ratio of the difference to the standard performance data as the first deviation degree. Taking hardware event A as an example, the difference is 5% and the standard performance data is 90%, then the first deviation degree is 5% / 90% ≈ 0.056. The first deviation degree is expressed in the form of a ratio, which can more intuitively reflect the relative gap between the real-time performance data and the standard performance data, facilitating subsequent performance bottleneck judgment and optimization decision-making. It should be emphasized that the calculation method of the second deviation degree is similar to that of the first deviation degree, which will not be elaborated here.

[0049] Figure 2 shows the flowchart of a method for optimizing the hardware performance of a processor according to an embodiment of the present disclosure Figure 2 , as Figure 2 shown, a method for optimizing the hardware performance of a processor includes: Step S201, in response to the activation of the performance monitoring function for the processor, control all target hardware event controllers on the processor to monitor the real-time performance data of the corresponding hardware events.

[0050] Step S202, determine the first deviation degree between the real-time performance data and the corresponding standard performance data.

[0051] Step S203, bind the hardware events corresponding to the first deviation degree that meets the first threshold to the hardware event counter, and monitor the target number of times of the corresponding hardware events based on the hardware event counter.

[0052] Step S204, determine the performance bottleneck of the processor based on the target number of times.

[0053] The specific implementation details of steps S201 - S204 are similar to those of steps S101 - S104, which will not be elaborated here.

[0054] Step S205, in response to the performance bottleneck existing in the core microarchitecture of the processor, monitor the processing duration of each processing stage in the instruction processing flow of the core microarchitecture.

[0055] In this embodiment, when it is determined that a performance bottleneck exists in the core microarchitecture of the processor, the instruction processing flow of the core microarchitecture is further monitored. The instruction processing flow of the core microarchitecture generally includes stages such as instruction fetching, decoding, execution, memory access, and write-back. By monitoring the processing duration of each processing stage, the performance of the core microarchitecture in different stages can be analyzed in detail. For example, monitor the interval duration between two instruction fetches in the instruction fetching stage, the interval duration between two decodings in the decoding stage, the duration for each execution unit to complete the instruction data operation in the execution stage, the duration of the memory access operation for each instruction in the memory access stage, and the duration for each instruction to be written back to the target storage location in the processor in the write-back stage. These processing duration data provide a detailed performance analysis basis for the optimization of the core microarchitecture. Among them. The instruction fetching stage is used to read instructions from memory according to the value of the Program Counter (PC); the decoding stage is used to send the instructions in the instruction register to the decoder for decoding, and the decoder converts the instructions into a series of control signals, which are used to control other components of the CPU to perform corresponding operations; the execution stage is used to make the various components of the CPU work together according to the control signals generated in the decoding stage to perform the operations specified by the instructions; the memory access stage is used to perform memory access operations if the instructions need to read data from memory or write data to memory; the write-back stage is used to write the results of the instruction execution back to the registers or memory of the CPU. If it is an arithmetic operation instruction, the result will be written back to the destination register. If it is a data transfer instruction, the data will be written to memory or other destinations.

[0056] Step S206, perform corresponding optimization measures on the processing stage whose processing duration meets the third threshold.

[0057] In this embodiment, the third threshold is a preset value used to determine whether the processing duration of a processing stage is abnormal. When the processing duration of a certain processing stage exceeds the third threshold, it indicates that this stage may be the location of the performance bottleneck of the core microarchitecture. Corresponding optimization measures are taken according to different processing stages. For example, if the processing durations of the instruction fetch stage and the instruction decode stage exceed the third threshold, the instruction prefetch function can be enabled, the instruction issue width can be increased, and / or the mapping relationship of the Instruction Translation Lookaside Buffer (ITLB) can be optimized; if the processing durations of the execution stage and the memory access stage exceed the third threshold, the instruction execution width can be increased and / or the mapping relationship of the Data Translation Lookaside Buffer (DTLB) can be optimized; if the processing duration of the write-back stage exceeds the third threshold, the cache depth of the ReOrder Buffer (ROB) can be increased and / or fast processing of the write-back state can be performed. By optimizing the performance bottleneck of the processing stage, the performance of the core microarchitecture can be effectively improved, thereby improving the overall performance of the processor.

[0058] In another embodiment, "monitoring the processing duration of each processing stage in the instruction processing flow of the core microarchitecture" in step S205 includes: Monitoring the interval duration between two consecutive instruction fetches in the instruction fetch stage.

[0059] In this embodiment, in the instruction fetch stage of the core microarchitecture, by monitoring the interval duration between two consecutive instruction fetch operations, the instruction fetch interval duration can be determined based on the number of clock cycles (T_Cycle_fetch) between the two instruction fetches. The instruction fetch interval duration reflects the efficiency of the instruction fetch operation. If the instruction fetch interval duration is long, it may mean that there is a bottleneck in the instruction fetch operation, such as low instruction cache hit rate, insufficient instruction prefetch, etc. By monitoring the instruction fetch interval duration, the performance problems in the instruction fetch stage can be detected in a timely manner, providing a basis for subsequent optimization measures.

[0060] Monitoring the interval duration between two consecutive instruction decodes in the instruction decode stage.

[0061] In this embodiment, in the instruction decode stage, by monitoring the interval duration between two consecutive instruction decode operations, the interval duration between the two instruction decode operations can be determined by the number of clock cycles (T_Cycle_decode) between the two instruction decode operations. The decode interval duration reflects the speed of instruction decoding. If the decode interval duration is long, it may indicate that the decoder is not efficient, such as high instruction complexity, unreasonable decoder design, etc. By monitoring the decode interval duration, the performance of the instruction decode stage can be accurately evaluated, providing a reference for further performance optimization.

[0062] Monitor the time taken for each execution unit to complete the instruction data operation during the execution phase.

[0063] In this embodiment, during the execution phase, to monitor the time taken for each execution unit to complete the instruction data operation, the time taken to complete the instruction data operation can be determined by the clock cycles for each execution unit to complete the instruction data operation. The execution time reflects the operation efficiency of the execution unit. If the execution time of a certain execution unit is relatively long, it may indicate that there are performance bottlenecks in this execution unit, such as excessive load on the execution unit or complex instruction dependency relationships. By monitoring the execution time, the location of the performance bottleneck in the execution phase can be determined, providing data support for optimizing the performance of the execution unit.

[0064] Monitor the time taken for the memory access operation of each instruction during the memory access phase.

[0065] In this embodiment, during the memory access phase, to monitor the time taken for the memory access operation of each instruction, the time taken for the memory access operation can be determined by the clock cycles occupied by each instruction's memory access operation. The memory access time reflects the efficiency of the memory access operation. If the memory access time is relatively long, it may indicate that there are bottlenecks in the memory access operation, such as low cache hit rate or high memory latency. By monitoring the memory access time, the performance problems in the memory access phase can be detected in a timely manner, providing a basis for optimizing the memory access operation.

[0066] Monitor the time taken for each instruction to be written back to the target storage location in the processor during the write-back phase.

[0067] In this embodiment, during the write-back phase, to monitor the time taken for each instruction to be written back to the target storage location in the processor, the time taken for each instruction to be written back to the target storage location in the processor (such as a register or memory) can be determined by the clock cycles occupied by the instruction being written back to the target storage location in the processor. The write-back time reflects the efficiency of the write-back operation. If the write-back time is relatively long, it may indicate that there are bottlenecks in the write-back operation, such as an overly long write-back queue or insufficient write-back bandwidth. By monitoring the write-back time, the performance of the write-back phase can be accurately evaluated, providing a reference for further performance optimization.

[0068] In another embodiment, step S206 "Execute the corresponding optimization measures for the processing phase whose processing time meets the third threshold" includes: In response to the processing time of the instruction fetch phase and the decoding phase meeting the third threshold, enable the instruction prefetch function, increase the instruction issue width, and / or optimize the mapping relationship of the instruction translation lookaside buffer (ITLB); the ITLB stores the mapping relationship between the virtual address and the physical address corresponding to the instruction.

[0069] In this embodiment, when the processing durations of the instruction fetch stage and the instruction decode stage exceed a third threshold, it indicates that there may be performance bottlenecks in these two stages. At this time, the instruction prefetch function can be enabled to load subsequent instructions into the instruction cache in advance, reducing the instruction fetch waiting time; increase the instruction issue width to improve the parallel processing ability of instructions; optimize the mapping relationship of the instruction translation lookaside buffer (ITLB) to improve the efficiency of instruction address translation. Among them, the processing durations of the instruction fetch stage and the instruction decode stage meeting the third threshold can mean that the ratio of the processing duration of the instruction fetch stage to the total duration of all processing stages in the instruction processing flow is greater than a certain threshold, the ratio of the processing duration of the instruction decode stage to the total duration of all processing stages in the instruction processing flow is greater than a certain threshold, or the ratio of the sum of the processing durations of the instruction fetch stage and the instruction decode stage to the total duration of all processing stages in the instruction processing flow is greater than a certain threshold.

[0070] In response to the processing durations of the execution stage and the memory access stage meeting the third threshold, increase the instruction execution width and / or optimize the mapping relationship of the data translation lookaside buffer (DTLB); the DTLB stores the mapping relationship between the virtual address and the physical address corresponding to the data.

[0071] In this embodiment, when the processing durations of the execution stage and the memory access stage exceed a third threshold, it indicates that there may be performance bottlenecks in these two stages. At this time, the instruction execution width can be increased to improve the parallel processing ability of the execution unit; optimize the mapping relationship of the data translation lookaside buffer (DTLB) to improve the efficiency of data address translation. These optimization measures can effectively improve the performance of the execution and memory access stages, thereby improving the overall performance of the core microarchitecture. Among them, the processing durations of the execution stage and the memory access stage meeting the third threshold can mean that the ratio of the processing duration of the execution stage to the total duration of all processing stages in the instruction processing flow is greater than a certain threshold, the ratio of the processing duration of the memory access stage to the total duration of all processing stages in the instruction processing flow is greater than a certain threshold, or the ratio of the sum of the processing durations of the execution stage and the memory access stage to the total duration of all processing stages in the instruction processing flow is greater than a certain threshold.

[0072] In response to the processing duration of the write-back stage meeting the third threshold, increase the cache depth of the reorder buffer (ROB) and / or perform fast processing of the write-back state; the fast processing of the write-back state includes early write-back, enhanced register renaming, and / or increased write-back bandwidth.

[0073] In this embodiment, when the processing duration of the write-back stage exceeds the third threshold, it indicates that there may be a performance bottleneck in the write-back stage. At this time, the cache depth of the reorder buffer (ROB) can be increased to improve the buffering capacity of the write-back operation; perform fast processing of the write-back status, such as early write-back, enhanced register renaming, and / or increased write-back bandwidth. These optimization measures can effectively improve the performance of the write-back stage, thereby enhancing the overall performance of the core microarchitecture.

[0074] Since RISC-V CPUs are increasingly applied to professional computing, scientific computing, and other fields, one characteristic of these fields is that their program architectures are relatively fixed, that is, the software programs running on RISC-V CPUs are relatively fixed. However, the core microarchitecture of RISC-V CPUs may not necessarily achieve the best performance under the current software architecture. Through the present disclosure, the core microarchitecture of RISC-V CPUs can be adjusted to the optimal performance state.

[0075] Figure 3 The flowchart shows a hardware optimization method for processor performance according to an embodiment of the present disclosure Figure 3 , as Figure 3 shown, a hardware optimization method for processor performance includes: Step S301, in response to the activation of the performance monitoring function for the processor, control all target hardware event controllers on the processor to monitor the real-time performance data of the corresponding hardware events.

[0076] Step S302, determine the first deviation degree between the real-time performance data and the corresponding standard performance data.

[0077] Step S303, bind the hardware events corresponding to the first deviation degree that meets the first threshold to the hardware event counter, and monitor the target number of times of the corresponding hardware events based on the hardware event counter.

[0078] Step S304, determine the performance bottleneck of the processor based on the target number of times.

[0079] The specific implementation details of steps S301 - S304 are similar to those of steps S101 - S104 and will not be elaborated here.

[0080] Step S305, in response to the performance bottleneck existing in the cache architecture of the processor, monitor the data hit rates of different cache parts in the cache architecture.

[0081] In this embodiment, when it is determined that a performance bottleneck exists in the cache architecture of the processor, the data hit rates of different cache parts in the cache architecture are further monitored. The cache parts generally include a translation lookaside buffer (TLB), a level 1 cache (L1 Cache), and a level 2 cache (L2 Cache). By monitoring the hit rates of the TLB, the L1 Cache, and the L2 Cache, the performance of the cache architecture at different levels can be analyzed in detail. For example, monitor the hit rate of the virtual address sent by the software to be accessed in the TLB to be able to normally query the physical address in the TLB; monitor the hit rate of accessing the data in the L1 Cache sent by the system bus; monitor the hit rate of accessing the data in the L2 Cache sent by the system bus. These hit rate data provide a detailed basis for performance analysis for subsequent optimization of the cache architecture.

[0082] This embodiment can be implemented by a cache architecture controller. Figure 6 The structure diagram of the cache architecture controller according to an embodiment of the present disclosure is shown. As Figure 6 shown, the cache architecture controller includes Hit_Detect_0, Hit_Detect_1, Hit_Detect_2, and Data_Prefetch. Among them, Hit_Detect_0, that is, the first access hit monitoring module, is used to monitor the hit rate of the virtual address sent by the software to be accessed in the TLB to be able to query the corresponding physical address in the TLB. When Hit_Detect_0 monitors that the hit rate of the virtual address accessed by the software in the TLB is lower than a certain threshold, such as 80%, an alarm signal is sent to Data_Prefetch, that is, the data prefetch module; Hit_Detect_1, that is, the second access hit monitoring module, whose function is to monitor the hit rate hit_L1 of accessing the data in the L1 Cache sent by the system bus, that is, the probability that the data to be accessed sent by the software can be queried in the local L1 Cache. When it is monitored that hit_L1 is less than a certain threshold, such as 75%. An alarm signal is sent to Data_Prefetch; Hit_Detect_2, that is, the third access hit monitoring module, whose function is to monitor the hit rate hit_L2 of accessing the data in the L2 Cache sent by the system bus, that is, the probability that the data to be accessed sent by the software can be queried in the local L2 Cache. When it is monitored that hit_L2 is less than a certain threshold, such as 70%. An alarm signal is sent to Data_Prefetch.

[0083] Step S306, perform corresponding optimization measures on the cache part whose data hit rate meets the fourth threshold.

[0084] In this embodiment, this step can be performed by Figure 6In the implementation of Data_Prefetch, the fourth threshold is a preset value used to determine whether the data hit rate of a certain cache part is abnormal. When the data hit rate of a certain cache part is lower than the fourth threshold, it indicates that there may be a performance bottleneck in that cache part. Corresponding optimization measures are taken according to different cache parts. For example, if the data hit rate of the TLB is lower than the fourth threshold, the page table storing the mapping relationship between virtual addresses and physical addresses in the double data rate (DDR) memory can be updated to the TLB based on the rules of software sending virtual addresses (such as consecutive access or jump access with a fixed step size, etc.); if the data hit rate of the L1 Cache is lower than the fourth threshold, the corresponding access addresses in the L2 Cache or DDR memory can be updated to the L1 Cache based on the rules of the system bus sending addresses (such as consecutive access or jump access with a fixed step size); if the data hit rate of the L2 Cache is lower than the fourth threshold, the corresponding access addresses in the DDR memory can be updated to the L2 Cache based on the rules of the system bus sending addresses. By optimizing the performance bottleneck of the cache part, the performance of the cache architecture can be effectively improved, thereby improving the overall performance of the processor.

[0085] In another embodiment, step S306 "executing corresponding optimization measures on the cache part whose data hit rate meets the fourth threshold" includes: In response to the data hit rate of the TLB meeting the fourth threshold, based on the rules of software sending virtual addresses, the page table storing the mapping relationship between virtual addresses and physical addresses in the double data rate DDR memory is updated to the TLB.

[0086] In this embodiment, when the data hit rate of the TLB is lower than the fourth threshold, it indicates that there may be a performance bottleneck in the TLB. At this time, according to the rules of software sending virtual addresses, the page table storing the mapping relationship between virtual addresses and physical addresses in the DDR memory is updated to the TLB. For example, if the software frequently accesses certain virtual addresses and the hit rate of these addresses in the TLB is low, the page table entries corresponding to these addresses can be loaded from the DDR memory to the TLB to increase the hit rate of the TLB, reduce the number of memory accesses, and thus improve the performance of the cache architecture.

[0087] In response to the data hit rate of the first-level cache meeting the fourth threshold, based on the rules of the system bus sending addresses, the corresponding access addresses in the second-level cache or DDR memory are updated to the first-level cache.

[0088] In this embodiment, when the data hit rate of the L1 Cache is lower than the fourth threshold, it indicates that there may be a bottleneck in the performance of the L1 Cache. At this time, according to the rule of sending addresses by the system bus, the corresponding access addresses in the L2 Cache or DDR memory are updated to the L1 Cache. For example, if the system bus frequently accesses certain addresses and the hit rate of these addresses in the L1 Cache is low, the data corresponding to these addresses can be prefetched from the L2 Cache to the L1 Cache. If these addresses are not in the L2 Cache, these addresses can be prefetched from the DDR memory to the L1 Cache, and at the same time, these addresses are stored in the L2 Cache to maintain data consistency. Thus, the hit rate of the L1 Cache can be increased, the number of accesses to the L2 Cache or DDR memory can be reduced, and the performance of the cache architecture can be improved.

[0089] In response to the data hit rate of the secondary cache meeting the fourth threshold, based on the rule of sending addresses by the system bus, the corresponding access addresses in the DDR memory are updated to the secondary cache.

[0090] In this embodiment, when the data hit rate of the L2 Cache is lower than the fourth threshold, it indicates that there may be a bottleneck in the performance of the L2 Cache. At this time, according to the rule of sending addresses by the system bus, the corresponding access addresses in the DDR memory are updated to the L2 Cache. For example, if the system bus frequently accesses certain addresses and the hit rate of these addresses in the L2 Cache is low, the data corresponding to these addresses can be prefetched from the DDR memory to the L2 Cache, improving the hit rate of the L2 Cache and reducing the number of accesses to the DDR memory, thereby improving the performance of the cache architecture.

[0091] Figure 5 FIG. shows a schematic structural diagram of a processor performance hardware optimization system according to an embodiment of the present disclosure, as Figure 5 shown, a processor performance hardware optimization system includes: A bottleneck analysis controller (Bottleneck_Analysis_Ctrl), configured to control all target hardware event controllers on the processor to monitor the real-time performance data of the corresponding hardware events in response to the activation of the performance monitoring function for the processor; determine the first deviation degree between the real-time performance data and the corresponding standard performance data; bind the hardware events corresponding to the first deviation degree that meets the first threshold to the hardware event counter, and monitor the target number of times of the corresponding hardware events based on the hardware event counter; determine the performance bottleneck of the processor based on the target number of times. A performance optimizer, configured to perform performance optimization on the processor based on the performance bottleneck.

[0092] In an implementable embodiment, the bottleneck analysis controller is further configured to: run the processor based on a standard test model program to obtain the original performance criteria corresponding to all hardware events on the processor; determine the second deviation degree between the original performance criteria and the standard performance data; in response to the second deviation degree meeting a second threshold, generate warning data indicating abnormal standard performance data and output it to the client; obtain new standard performance data; the new standard performance data is the standard performance data regenerated by the client based on the warning data, or the standard performance data obtained by weighted averaging the original performance criteria and the standard performance data.

[0093] In an implementable embodiment, the bottleneck analysis controller is further configured to: determine the difference between the real-time performance data and the corresponding standard performance data; determine the ratio of the difference to the standard performance data as the first deviation degree.

[0094] In an implementable embodiment, the performance optimizer includes a microarchitecture controller (Micro_Ctrl), and the microarchitecture controller is configured to: in response to a performance bottleneck existing in the core microarchitecture of the processor, monitor the processing duration of each processing stage in the instruction processing flow of the core microarchitecture; perform corresponding optimization measures on the processing stage whose processing duration meets a third threshold.

[0095] In an implementable embodiment, the microarchitecture controller is further configured to: monitor the interval duration between two consecutive instruction fetches in the instruction fetch stage; monitor the interval duration between two consecutive instruction decodings in the instruction decoding stage; monitor the duration for each execution unit to complete instruction data operations in the execution stage; monitor the duration of the memory access operation for each instruction in the memory access stage; monitor the duration for each instruction to be written back to the target storage location of the processor in the write-back stage.

[0096] In an implementable embodiment, the microarchitecture controller is further configured to: in response to the processing durations of the instruction fetch stage and the instruction decoding stage meeting the third threshold, enable the instruction prefetch function, increase the instruction issue width, and / or optimize the mapping relationship of the instruction translation lookaside buffer (ITLB); the ITLB stores the mapping relationship between the virtual address and the physical address corresponding to the instruction; in response to the processing durations of the execution stage and the memory access stage meeting the third threshold, increase the instruction execution width and / or optimize the mapping relationship of the data translation lookaside buffer (DTLB); the DTLB stores the mapping relationship between the virtual address and the physical address corresponding to the data; in response to the processing duration of the write-back stage meeting the third threshold, increase the cache depth of the reorder buffer (ROB) and / or perform fast write-back state processing; the fast write-back state processing includes early write-back, enhanced register renaming, and / or increased write-back bandwidth.

[0097] In one implementable manner, the performance optimizer includes a cache architecture controller (Cache_Ctrl), and the cache architecture controller is configured to: in response to a performance bottleneck existing in the cache architecture of the processor, monitor the data hit rates corresponding to different cache parts in the cache architecture; the cache parts include a translation lookaside buffer (TLB), a level-1 cache, and a level-2 cache; and perform corresponding optimization measures on the cache parts whose data hit rates meet a fourth threshold.

[0098] In one implementable manner, the cache architecture controller is further configured to: in response to the data hit rate of the TLB meeting the fourth threshold, update the page table storing the mapping relationship between the virtual address and the physical address in the double data rate (DDR) memory to the TLB based on the rule of the software sending the virtual address; in response to the data hit rate of the level-1 cache meeting the fourth threshold, update the corresponding access address in the level-2 cache or the DDR memory to the level-1 cache based on the rule of the system bus sending the address; and in response to the data hit rate of the level-2 cache meeting the fourth threshold, update the corresponding access address in the DDR memory to the level-2 cache based on the rule of the system bus sending the address.

[0099] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution of this disclosure can be achieved, and this is not limited herein.

[0100] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this disclosure, "a plurality of" means two or more unless otherwise specifically defined.

[0101] As described above, the above are only specific implementation manners of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by this disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.

Claims

1. A hardware optimization method for processor performance, characterized in that The method includes: In response to the activation of the performance monitoring function for the processor, controlling all target hardware event controllers on the processor to monitor the real-time performance data of corresponding hardware events; Determining a first deviation degree between the real-time performance data and corresponding standard performance data; Binding the hardware events corresponding to the first deviation degree that meets a first threshold to a hardware event counter, and monitoring the target number of corresponding hardware events based on the hardware event counter; Determining the performance bottleneck of the processor based on the target number; Performing performance optimization on the processor based on the performance bottleneck.

2. The method according to claim 1, characterized in that Before determining the first deviation degree between the real-time performance data and corresponding standard performance data, the method further includes: Running the processor based on a standard test model program to obtain the original performance standards corresponding to all hardware events on the processor; Determining a second deviation degree between the original performance standards and the standard performance data; In response to the second deviation degree meeting a second threshold, generating warning data for abnormal standard performance data and outputting it to the client; Obtaining new standard performance data; the new standard performance data is the standard performance data regenerated by the client based on the warning data, or the standard performance data obtained by weighted averaging the original performance standards and the standard performance data.

3. The method according to claim 1, wherein The determining of the first deviation degree between the real-time performance data and corresponding standard performance data includes: Determining the difference between the real-time performance data and corresponding standard performance data; Determining the ratio of the difference to the standard performance data as the first deviation degree.

4. The method according to claim 1, characterized in that, The performing of performance optimization on the processor based on the performance bottleneck includes: In response to the performance bottleneck existing in the core microarchitecture of the processor, monitoring the processing duration of each processing stage in the instruction processing flow of the core microarchitecture; Performing corresponding optimization measures on the processing stage whose processing duration meets a third threshold.

5. The method according to claim 4, characterized in that The monitoring of the processing duration of each processing stage in the instruction processing flow of the core microarchitecture includes: Monitoring the interval duration between two instruction fetches in the instruction fetch stage; Monitoring the interval duration between two instruction decodings in the instruction decoding stage; Monitoring the duration for each execution unit to complete instruction data operations in the execution stage; Monitoring the duration of the memory access operation for each instruction in the memory access stage; Monitoring the duration for each instruction to be written back to the target storage location of the processor in the write-back stage.

6. The method according to claim 5, characterized in that The performing of corresponding optimization measures on the processing stage whose processing duration meets a third threshold includes: In response to the processing durations of the instruction fetch stage and the instruction decoding stage meeting the third threshold, activating the instruction prefetch function, increasing the instruction issue width, and / or optimizing the mapping relationship of the instruction translation lookaside buffer (ITLB); the ITLB stores the mapping relationship between the virtual address and the physical address corresponding to the instruction; In response to the processing durations of the execution stage and the memory access stage meeting the third threshold, increasing the instruction execution width and / or optimizing the mapping relationship of the data translation lookaside buffer (DTLB); the DTLB stores the mapping relationship between the virtual address and the physical address corresponding to the data; In response to the processing duration of the write-back stage satisfying a third threshold, increase the cache depth of the reorder buffer ROB and / or perform fast processing of the write-back status; the fast processing of the write-back status includes early write-back, enhanced register renaming, and / or increased write-back bandwidth.

7. The method according to claim 1, wherein The performance optimization of the processor based on the performance bottleneck includes: In response to the performance bottleneck existing in the cache architecture of the processor, monitor the data hit rates corresponding to different cache parts in the cache architecture; the cache parts include the translation lookaside buffer TLB, the first-level cache, and the second-level cache; Execute corresponding optimization measures on the cache part whose data hit rate satisfies a fourth threshold.

8. The method according to claim 7, wherein The execution of corresponding optimization measures on the cache part whose data hit rate satisfies a fourth threshold includes: In response to the data hit rate of the TLB satisfying the fourth threshold, update the page table storing the mapping relationship between virtual addresses and physical addresses in the double data rate DDR memory to the TLB based on the rules for the software to send virtual addresses; In response to the data hit rate of the first-level cache satisfying the fourth threshold, update the corresponding access address in the second-level cache or the DDR memory to the first-level cache based on the rules for the system bus to send addresses; In response to the data hit rate of the second-level cache satisfying the fourth threshold, update the corresponding access address in the DDR memory to the second-level cache based on the rules for the system bus to send addresses.

9. A processor performance hardware optimization system, characterized in that, The system includes: A bottleneck analysis controller, configured to, in response to the activation of the performance monitoring function for the processor, control all target hardware event controllers on the processor to monitor the real-time performance data of corresponding hardware events; determine the first deviation degree between the real-time performance data and the corresponding standard performance data; bind the hardware events corresponding to the first deviation degree that satisfies a first threshold to a hardware event counter, and monitor the target number of corresponding hardware events based on the hardware event counter; determine the performance bottleneck of the processor based on the target number; A performance optimizer, configured to perform performance optimization on the processor based on the performance bottleneck.

10. The system according to claim 9, characterized in that The bottleneck analysis controller is further configured to: Run the processor based on a standard test model program to obtain the original performance standards corresponding to all hardware events on the processor; Determine the second deviation degree between the original performance standards and the standard performance data; In response to the second deviation degree satisfying a second threshold, generate warning data for the abnormal standard performance data and output it to the user side; Obtain new standard performance data; the new standard performance data is the standard performance data regenerated by the user side based on the warning data, or the standard performance data obtained by performing weighted averaging on the original performance standards and the standard performance data.

11. The system according to claim 9, characterized in that, The bottleneck analysis controller is further configured to: Determine the difference between the real-time performance data and the corresponding standard performance data; Determine the ratio of the difference to the standard performance data as the first deviation degree.

12. The system according to claim 9, wherein The performance optimizer includes a microarchitecture controller, and the microarchitecture controller is configured to: In response to the presence of the performance bottleneck in the core microarchitecture of the processor, monitor the processing duration of each processing stage in the instruction processing flow of the core microarchitecture; Execute corresponding optimization measures for the processing stage whose processing duration meets the third threshold.

13. The system according to claim 12, characterized in that, The microarchitecture controller is further configured to: Monitor the interval duration between two consecutive instruction fetches in the instruction fetch stage; Monitor the interval duration between two consecutive decodings in the decoding stage; Monitor the duration for each execution unit to complete the instruction data operation in the execution stage; Monitor the duration of the memory access operation for each instruction in the memory access stage; Monitor the duration for each instruction to be written back to the target storage location of the processor in the write-back stage.

14. An electronic device, characterized in that, including: The processor performance hardware optimization system according to any one of claims 9-13.

15. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Method and device for monitoring system performance

    CN104239183A

  • Method and system for testing performance

    CN106855844A

  • Application program performance data processing method and device and storage medium

    CN110362460A

  • Method for monitoring pipeline instruction execution

    CN115454505A

  • Performance monitoring system and method and electronic equipment

    CN116166518A