Chip optimization processing method and system, and chip

By acquiring behavior prediction rules from the on-chip system and using analysis software to optimize the operation of the on-chip system, the problem of hardware accelerators being unable to cover other software and being costly is solved, thus achieving efficient performance optimization of the processor.

WO2026158326A1PCT designated stage Publication Date: 2026-07-30BEIJING YUNYAO XINDAO TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING YUNYAO XINDAO TECHNOLOGY CO LTD
Filing Date
2026-01-20
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Traditional processors, when faced with specific software, use hardware accelerators, which can lead to problems such as the inability to cover other software and high costs.

Method used

By acquiring behavior prediction rules from the on-chip system, analyzing and generating rules based on historical operating information, measures can be taken in advance to optimize the operation of the on-chip system. Analysis software can be used to analyze operating information in external storage space and generate rules to guide the future operation of the on-chip system.

Benefits of technology

It achieves processor performance optimization for different target programs, avoiding increased hardware costs and improving processor performance and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2026073815_30072026_PF_FP_ABST
    Figure CN2026073815_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a chip optimization processing method and system, and a chip. The method comprises: a system-on-chip located on a chip acquiring a rule, wherein the rule reflects a behavior prediction when the system-on-chip runs target software, the rule is obtained by means of analysis on the basis of running information collected by means of the system-on-chip when historically running the target software, and the running information is generated by means of the system-on-chip by running the target software; and when running the target software, the system-on-chip taking measures in advance according to the rule, wherein the measures taken in advance predict operations to be performed by the system-on-chip and prepare for the operations. By means of the present application, the problems of higher costs and inability to support other software caused by the use of a hardware accelerator for processor optimization of specific software are solved, such that processor performance can be optimized for different target programs.
Need to check novelty before this filing date? Find Prior Art

Description

A chip optimization processing method, system, and chip Technical Field

[0001] This application relates to the field of chips, and more specifically, to a chip optimization processing method, system, and chip. Background Technology

[0002] Traditional processors are primarily designed for general-purpose computing scenarios, aiming to achieve relatively balanced performance across a wide range of flexible programs. However, this inevitably leads to performance fluctuations due to different programs. In reality, some scenarios involve running the same program for extended periods, including but not limited to cloud computing and dedicated programs such as search, recommendation, advertising, and EDA running on servers for long periods.

[0003] To further improve software efficiency, the most obvious approach is to increase processor performance. However, in practical applications, designing dedicated hardware for a program is very expensive. Another approach is to use a Domain Specific Accelerator (DSA). A DSA is a hardware architecture specifically designed for a particular program to achieve optimal performance, but its drawbacks include limited coverage of diverse scenarios and high cost. Summary of the Invention

[0004] This application provides a chip optimization processing method, system, and chip to at least solve the problems of insufficient coverage of other software and high cost caused by using hardware accelerators when optimizing processors for specific software.

[0005] According to one aspect of this application, a chip optimization processing method is provided, comprising: obtaining rules for an on-chip system located on the chip, wherein the rules reflect a prediction of the behavior of the on-chip system when running target software, the rules being obtained by analyzing operational information collected from the on-chip system's historical execution of the target software, the operational information being generated by the on-chip system running the target software; and the on-chip system taking pre-emptive measures according to the rules when running the target software, wherein the measures are determined according to the rules, the pre-emptive measures predicting future operations of the on-chip system and preparing for such operations.

[0006] Furthermore, it also includes: analysis software receiving the runtime information, wherein the runtime information is obtained through an interface pre-configured on the chip, and after being read from the interface, the runtime information is stored in a storage space located outside the chip; the analysis software analyzes the runtime information and determines the rule; the analysis software sends the rule to the system-on-a-chip; or, the system-on-a-chip obtaining the rule includes: the analysis software receiving the runtime information; the analysis software obtaining parameters for generating the rule based on the runtime information; the system-on-a-chip receiving the parameters and using the parameters in an algorithm running on the system-on-a-chip to generate the rule; wherein the algorithm is used to generate the rule.

[0007] Furthermore, the computing resources for running the analysis software are different from those for the on-chip system running the target software.

[0008] Furthermore, in the case where the chip includes multiple cores, each core comprises two parts: a first part for running the analysis software and a second part for running the target software as a system-on-a-chip; or, the analysis software runs on a computing device different from the computing device on which the chip is located; or, the analysis software runs on the system-on-a-chip, which uses different time-domain resources to run the analysis software and the target software in a time-sharing manner.

[0009] Furthermore, the target software runs on multiple chips, which are located on different computing devices; the analysis software receives runtime information generated by the on-chip system running the target software on each of the multiple chips; the analysis software analyzes the runtime information generated by all the on-chip systems to obtain the rules; the analysis software transmits the rules to the on-chip system on each of the multiple chips.

[0010] According to another aspect of this application, a chip is also provided, including a central processing unit and a cache; wherein, the system-on-a-chip on the chip is configured to perform the following method steps: obtaining rules, wherein the rules reflect a prediction of the behavior of the system-on-a-chip when running target software, the rules being obtained by analyzing runtime information collected from the historical runtime of the system-on-a-chip when running the target software, the runtime information being generated by the system-on-a-chip running the target software; and taking measures in advance according to the rules when running the target software, wherein the measures are determined according to the rules, the measures taken in advance predicting future operations of the system-on-a-chip and preparing for the operations.

[0011] Furthermore, it also includes: an interface, wherein there are one or more interfaces, the interfaces being used to send the operating information to an external part of the on-chip system, and to receive rules from the analysis software or to receive parameters from the analysis software; wherein the parameters are used in an algorithm running on the on-chip system to generate the rules, and the algorithm is used to generate the rules.

[0012] According to another aspect of this application, a chip optimization processing system is also provided, including the aforementioned chip and analysis software, wherein the analysis software is used to perform the following method steps: receiving running information, wherein the running information is obtained through an interface pre-configured on the chip, and after being read from the interface, the running information is stored in a storage space located outside the chip; analyzing the running information to determine the rule; sending the rule to the on-chip system; or, obtaining parameters for generating the rule based on the running information, and sending the parameters to the on-chip system; wherein an algorithm running on the on-chip system uses the parameters to generate the rule; the algorithm is used to generate the rule.

[0013] Furthermore, the computing resources for running the analysis software are different from those for the on-chip system running the target software.

[0014] Furthermore, in the case where the chip includes multiple cores, each core comprises two parts: a first part for running the analysis software and a second part for running the target software as a system-on-a-chip; or, the analysis software runs on a computing device different from the computing device on which the chip is located; or, the analysis software runs on the system-on-a-chip, which uses different time-domain resources to run the analysis software and the target software in a time-sharing manner.

[0015] In this embodiment, on-chip system-on-a-chip (SoC) acquisition rules are employed. These rules reflect predictions of the SoC's behavior when running target software. The rules are derived from analysis of historical runtime information collected during the SoC's operation of the target software. This runtime information is generated by the SoC running the target software. When running the target software, the SoC takes pre-emptive measures based on these rules. These measures are determined according to the rules, and the pre-emptive measures predict and prepare for future operations by the SoC. This application solves the problems of insufficient coverage of other software and high costs associated with using hardware accelerators for processor optimization of specific software, thereby enabling processor performance optimization for different target programs. Attached Figure Description

[0016] The accompanying drawings, which form part of this application, are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an undue limitation thereof. In the drawings:

[0017] Figure 1 is a schematic diagram of computer hardware and software architecture based on related technologies;

[0018] Figure 2 is a schematic diagram of target software operation optimization according to an embodiment of this application;

[0019] Figure 3 is a second schematic diagram of target software operation optimization according to an embodiment of this application;

[0020] Figure 4 is a schematic diagram of the target software operation optimization according to an embodiment of this application;

[0021] Figure 5 is a schematic diagram of the interaction timing according to an embodiment of this application;

[0022] Figure 6 is a second interactive timing diagram according to an embodiment of this application;

[0023] Figure 7 is a schematic diagram of the interaction timing according to an embodiment of this application;

[0024] Figure 8 is an interactive flowchart according to an embodiment of this application;

[0025] Figure 9 is a schematic diagram of the interaction flow according to an embodiment of this application; and,

[0026] Figure 10 is a flowchart of a chip optimization processing method according to an embodiment of this application. Detailed Implementation

[0027] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0028] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0029] In related technologies, processors typically use more resources to improve performance. For example, to increase cache hit rate, traditional processor manufacturers mainly increase cache area and cached data volume to improve performance. This approach increases cost and processor size. The following implementation focuses on fully processing and analyzing the historical execution information of the target software or program to predict future execution trajectories. This method avoids trading area for performance; instead, it improves performance through prediction while saving space.

[0030] The technical terms used in the following embodiments will be explained first.

[0031] System-on-a-Chip

[0032] In the following embodiments, the system within the chip (or processor) is referred to as a system-on-a-chip (SoC). An SoC may include a central processing unit (CPU), cache, etc. In some literature, the CPU can be understood as an SoC; in others, it can be understood as the central computing unit on an SoC. To avoid this ambiguity, the following embodiments use both SoC and central computing unit.

[0033] Offline software or offline programs

[0034] In the following embodiments, "online program" refers to a program that runs in real time. For example, target software is an online program that runs in real time for high performance. Analysis software is an offline program, which can run offline in a time-separated manner or on other machines. Analysis software can also be an online program in certain scenarios for high-performance online analysis. It should be noted that if multiple kernels exist, the kernels can be divided into two parts, one running online programs and the other running offline programs. In one example, processing running inside the chip is real-time hardware processing. However, if it runs in real-time on the on-chip system at the operating system level, it can also be considered an online program. Offline is more of a separation in terms of time and space; for example, programs that do not need to run in real time are called offline programs, programs that can run on other systems are called offline programs, and scenarios requiring high-performance execution will choose online implementation.

[0035] Figure 1 is a schematic diagram of computer hardware and software architecture according to related technologies. As shown in Figure 1, the application or program runs on the operating system and then runs on the chip through the compiler. In Figure 1, the application or program, the operating system and the compiler are all referred to as the software part, and the chip and the on-chip system running inside the chip are referred to as the hardware part.

[0036] i-cache and d-cache

[0037] The i-cache (instruction cache) and d-cache (data cache) play different roles to improve the execution efficiency and overall system performance of the on-chip system. The i-cache is a high-speed cache specifically designed to store instructions that the on-chip system is about to execute. When the on-chip system needs to execute an instruction, it first checks if the instruction is already cached in the i-cache. If the instruction is already cached in the i-cache (a cache hit), the on-chip system can directly read the instruction from the i-cache without waiting to fetch it from memory. This significantly improves instruction access speed, thereby accelerating program execution. The d-cache is a high-speed cache specifically designed to store recently used data on the on-chip system, such as variables and arrays. When the on-chip system needs to access this data, it checks if the data is already cached in the d-cache. If the data is already cached in the d-cache, the on-chip system can directly read it from the d-cache, avoiding the latency caused by frequent data reads from memory. This improves data access speed, thereby enhancing overall program performance.

[0038] In one example, a system-on-a-chip (SoC) may employ multi-level caching, where the i-cache and d-cache can be part of a multi-level caching architecture integrated within the SoC. For instance, a typical SoC might contain three levels of cache: L1, L2, and L3. The L1 cache, being closest to the SoC's core, is usually divided into an i-cache and a d-cache. When the SoC executes a program, it first searches for the required instructions and data in the L1 cache. If the required data or instructions are not found in the L1 cache, the SoC will then search in the L2 or L3 cache.

[0039] In the following implementation, in order to improve the performance of the system-on-chip, the operation of the system-on-chip is predicted, so that instructions and / or data can be prefetched in advance. This prefetching may include configuring future needed instructions in the i-cache and / or configuring future needed data in the d-cache.

[0040] Overview of pagination mechanism

[0041] Paging is a finer-grained memory space partitioning mechanism than segmentation. When paging is enabled, the processor must use the page management mechanism to translate linear addresses into physical addresses. The paging mechanism divides both the linear address space and the physical address space into fixed-size pages (usually 4KB) and maintains a page translation table structure through the Memory Management Unit (MMU) to complete the mapping and translation from linear address space pages to physical address pages.

[0042] Replacement algorithm

[0043] Central Processing Unit (CPU) replacement algorithms are algorithms used to manage CPU time slice allocation. They primarily determine which programs will be prioritized for allocation when multiple programs are running concurrently. These algorithms consider factors such as program priority, CPU load, I / O operation frequency, and size. Several replacement algorithms from related technologies are introduced below.

[0044] Optimal Page Replacement Algorithm: This is a theoretical algorithm that selects pages that will not be accessed for the longest period of time to replace them, thereby minimizing the page error rate. However, due to its complexity and high computational cost, it is rarely used in practical systems.

[0045] First-In, First-Out (FIFO) Algorithm: This algorithm replaces pages in the order they enter memory, with the earliest page being replaced first. This algorithm is simple to implement, but it can lead to the "Belady's phenomenon," where increasing the number of allocated pages actually increases the page fault rate.

[0046] Least Recently Used (LRU) algorithm: This algorithm selects the least recently used page for replacement. The LRU algorithm can reflect the page usage trend well, but it requires more management overhead.

[0047] Clock Algorithm: This is an approximation of the LRU algorithm, using the ticking of a simulated clock to determine which pages should be replaced. This algorithm reduces management overhead, but its performance is inferior to LRU in some situations.

[0048] The technical solutions described below can be applied to the replacement algorithm.

[0049] Translation Backup Buffer

[0050] The Translation Lookaside Buffer (TLB), also known as the page table cache or address-bypass cache, is an on-chip cache used by the memory management unit to improve the speed of virtual-to-physical address translation. All desktop and server chips (such as x86) use the TLB. The TLB has a fixed number of slots for storing tag page table entries that map virtual addresses to physical addresses. It is a typical example of content-addressable memory (CAM). Its search keyword is the virtual memory address, and its search result is the physical address. If the requested virtual address exists in the TLB, the CAM will provide a very fast match, and the obtained physical address can then be used to access memory. If the requested virtual address is not in the TLB, the tag page table is used for virtual-to-physical address translation, but accessing the tag page table is much slower than using the TLB. Some systems allow the tag page table to be swapped to secondary memory, in which case virtual-to-physical address translation can take a very long time.

[0051] The technical solutions described below can also be applied to TLB.

[0052] Target software

[0053] The target software is software that runs on top of an operating system. In one example, the target software runs online in real time. The analysis software can be offline, or online in high-performance scenarios. The target software can have various functions; for example, in a big data system, the target software can be software that performs data processing; in a video system, the target software can be software that performs image processing, etc. In the following embodiments, the target software can be long-running software; for example, license plate recognition and warning software in a video surveillance system needs to run for extended periods. Optimization measures as described in the following embodiments can be used for such software to make its operation more efficient. To distinguish it from the target software, in the following embodiments, the software that performs chip information analysis (the chip information offline processing program) is referred to as log analysis software (or simply analysis software). This analysis software can also be called the first software (in the following text, offline software refers to the first software), and the target software can also be called the second software.

[0054] In order to solve the problems in the related technology, a chip optimization processing method is provided in the following embodiments. FIG10 is a flowchart of the chip optimization processing method according to an embodiment of the present application. As shown in FIG10, the steps included in the method in FIG10 will be described below.

[0055] Step S102: The on-chip system on the chip acquires rules, wherein the rules reflect the behavior prediction of the on-chip system when running the target software, and the rules are obtained by analyzing the running information collected when the on-chip system runs the target software in the past, and the running information is generated by the on-chip system running the target software.

[0056] The rules used in this step can be analyzed in various ways. As an optional implementation, they can be obtained by analyzing the runtime information using analysis software. The analysis software can process the rules through a pre-configured step: the analysis software receives the runtime information, which is obtained through an interface pre-configured on the chip. After the runtime information is read from the interface, it is stored in a storage space (e.g., memory) outside the chip; the analysis software analyzes the runtime information and then determines the rules; the analysis software sends the rules to the system-on-a-chip.

[0057] Regarding the rules, it can be determined whether further optimization is needed based on the operating status of the on-chip system. Specifically, as an optional implementation, after obtaining the rules based on the operating information, the on-chip system runs the target software according to the rules, recording performance data during this process. If the performance data meets expectations, the rules are deemed compliant. If the performance data does not meet expectations, the on-chip system continues to run the target software to collect operating information. The analysis software then regenerates the rules based on all collected operating information. The on-chip system runs the target software according to the regenerated rules, records performance data again, and determines whether it meets expectations. This process is repeated until the performance data meets expectations.

[0058] Step S104: When the on-chip system is running the target software, it takes measures in advance according to the rules, wherein the measures are determined according to the rules, and the measures taken in advance estimate the operations that the on-chip system will perform in the future and prepare for the operations.

[0059] Step S106: Different rules are used when the on-chip system runs different target software.

[0060] The above steps solve the problems of not being able to cover other software and high costs caused by using hardware accelerators when optimizing processors for specific software, thus enabling processor performance optimization for different target programs.

[0061] In the above steps, the on-chip system can take many measures according to the rules, which will be illustrated by examples below.

[0062] Example 1: The on-chip system determines the cache to be used in the future when running the target software according to the rules, and prepares the cache to be used in the future in advance. The cache includes at least one of the following: instruction cache and data cache.

[0063] Example 2: The on-chip system determines the instruction branches to be used in the future when running the target software according to the rules, and prepares for the instruction branches to be used in the future in advance.

[0064] Example 3: The on-chip system determines the translation backup buffer to be used in the future when running the target software according to the rules, and prepares the translation backup buffer to be used in the future in advance.

[0065] It should be noted that the above three examples can be used individually or together; of course, other measures can also be taken.

[0066] For the analysis software, the computing resources running the analysis software and the on-chip system running the target software can be different. For example, if the chip includes multiple cores, each core may consist of two parts: a first part runs the analysis software, and a second part runs the target software as an on-chip system. Alternatively, the analysis software may run on a computing device that is different from the computing device on which the chip resides. Another example is that the target software runs on multiple chips, each located on a different computing device; the analysis software receives runtime information generated by the on-chip system running the target software on each of the multiple chips; the analysis software analyzes all the collected runtime information generated by the on-chip systems to obtain the rules; and the analysis software transmits the rules to the on-chip system on each of the multiple chips.

[0067] In another example, the system-on-a-chip (SoC) can run an algorithm that generates rules. In this case, the SoC acquiring the rules includes: the analysis software receiving the runtime information; the analysis software obtaining parameters for generating the rules based on the runtime information; the SoC receiving the parameters and using them in the algorithm running on the SoC to generate the rules; wherein the algorithm is used to generate the rules. In this example, the algorithm can be a neural network model, and the parameters can be the parameters of the neural network model or the weights of each layer.

[0068] The following explanation is based on the accompanying diagram.

[0069] Figure 2 is a schematic diagram of the target software operation optimization according to an embodiment of this application. The software side in Figure 2 corresponds to the software part in Figure 1, the hardware side in Figure 2 corresponds to the hardware part in Figure 1, the offline chip information processing program in Figure 2 corresponds to the analysis software described above, and the accelerated program running in real time in Figure 2 corresponds to the target software described above. The hardware side has multiple kernels, which can include two parts: the first part includes kernels 10, 11 to 1M, and the second part includes kernels 20, 21 to 2N. The second part of the kernel is used to run the accelerated program; this part is referred to as the program acceleration kernel. The first part of the kernel is used to run the analysis software; this part is referred to as the information processing kernel.

[0070] An interface is configured on the hardware side to retrieve runtime information (i.e., hardware-side chip runtime information) obtained by the kernel during the execution of the target program from the on-chip system. This retrieved runtime information can be stored in memory. This memory is not the on-chip cache; rather, it has a larger storage capacity and can hold a significant amount of runtime information. This runtime information will be used to predict the runtime behavior of the target software on the on-chip system, and the prediction results can guide the on-chip system in executing the target program.

[0071] The prediction work is performed by an offline chip information processing program (i.e., analysis software). (It should be noted that, in one example, in certain high-performance scenarios, online programs can also perform real-time analysis and real-time rule configuration; that is, the analysis software can be offline or online.) The analysis yields the characteristics of the on-chip system when running the target software. Through the analysis of these characteristics, rules (i.e., operating rules generated by the software side) are obtained. These rules predict the future behavior of the on-chip system when running the target software. After obtaining the rules, the rules are sent to the program acceleration kernel through memory. The program acceleration kernel processes the target software according to the rules.

[0072] In Figure 2, the analysis software is run by a portion of the kernel. In another implementation, it can also be executed by a separate host or computing resource. Figure 3 is a second schematic diagram of target software operation optimization according to an embodiment of this application. As shown in Figure 3, a separate host is used instead of the first portion of the kernel (i.e., the information processing kernel) in Figure 2. It should be noted that the separate host here refers to a chip that is not the same chip as the chip that runs the analysis software. This separate host can also be understood as a computing resource independent of the chip that runs the target software. This computing resource can be a single or clustered host, or it can be a cloud computing resource. In Figure 3, the target software (i.e., the accelerated program) runs through a chip. This chip can include multiple kernels, i.e., kernel 0, kernel 1 to kernel K in Figure 3, used to run the target software. During the operation, hardware-side chip operation information is collected. The software processes the program offline (i.e., the analysis software), generates operation rules based on the operation information (the operation rules here are the same as those in Figure 2), and sends the generated operation rules to the on-chip system on the chip. The on-chip system processes the target software according to the operation rules.

[0073] The target software may run on multiple different hosts or on different processors, where a processor can be understood as a chip. Figure 4 is a schematic diagram of target software operation optimization according to an embodiment of this application. As shown in Figure 4, the processors run in a distributed manner, with each processor running the target software. Each processor sends operation information to the analysis software. After summarizing the operation information, the analysis software analyzes the operation information to obtain operation rules (i.e., the guidance rules in Figure 4) and sends the guidance rules to the processors. It should be noted that the operation rules sent to each processor can be the same, or different guidance rules can be sent according to the attributes of the processor.

[0074] As can be seen from the above implementation methods, the software program can be divided into two parts: one part is the program to be accelerated that runs in real time (i.e., the target software or target program), and the other part is the software program that performs offline processing on the information generated by the real-time operation of the on-chip system (i.e., the analysis software or analysis program).

[0075] In Figures 2 to 4 above, the software side runs an offline analysis program and a real-time accelerated program, respectively. On the hardware side, the on-chip system's runtime information generated by the target program is transmitted to the offline analysis program on the software side. The offline analysis program can run on a portion of the information processing core of the same server, as shown in Figure 2, or it can run on other servers or PC hosts. The memory layer is an information exchange mechanism, which can be internally exchanged through memory as shown in Figure 2, or exchanged through a network or other means on other hosts as shown in Figure 3 or 4. In another example, the analysis software can also run on the on-chip system, which uses different time-domain resources to run the analysis software and the target software in a time-sharing manner. Regardless of how the analysis software and the target software use different computing resources, the analysis software only needs to obtain runtime information through the external storage device of the chip, then perform analysis, and finally optimize the operation of the target software on the on-chip system through rules. As an example, the external storage device can be memory outside the chip, and the chip is connected to this memory. The following explanation uses memory as an example.

[0076] Memory can be either off-chip or on-chip storage for the CPU. In addition to storing the instructions and data necessary for the program, it also stores key CPU information generated by the acceleration core (i.e., the kernel running the target program). The software uses this key information to calculate and generate new operating rules. These rules are then sent to the memory and configured to the program acceleration core (i.e., the kernel running the target program) to guide the CPU's operating behavior and improve system performance.

[0077] Figures 2 to 4 above illustrate several optional architectures. The following describes several optional hardware-software interaction timings. The system-on-a-chip (SoC) runs the target software during a first time period and sends the runtime information obtained during this first time period to the analysis software. The analysis software analyzes the runtime information during a second time period to obtain a first rule and transmits the first rule to the SoC. The SoC runs the target software according to the first rule during a third time period and sends the runtime information generated during this third time period to the analysis software. The analysis software analyzes all collected runtime information to generate a second rule and transmits it to the SoC, and so on. Alternatively, during the Nth time period, the SoC runs the target software according to the rule previously sent by the analysis software, and the analysis software analyzes the runtime information obtained previously during the Nth time period to obtain a new rule. During the N+1th time period, the SoC runs the target software according to the new rule, and the analysis software summarizes the runtime information from the Nth time period and analyzes it again to obtain a new rule, and so on. Alternatively, after the system-on-a-chip (SoC) has been running the target software for a period of time, the analysis software analyzes the running information obtained during that period to obtain rules and sends the rules to the SoC; thereafter, the SoC continues to use the rules.

[0078] The above timing sequences are explained below with reference to the accompanying diagrams.

[0079] Figure 5 is a timing diagram illustrating the interaction according to an embodiment of this application. Figure 5 shows the timing diagram for "intermittent processing." As shown in Figure 5, hardware operation refers to the on-chip system running the target software. During operation, it collects information generated by the target software (in Figure 5, this is hardware information or key information). After collecting the key information, the analysis software analyzes it to generate rules, and then sends the rules to the on-chip system for hardware operation. In Figure 5, the hardware operates using default rules within segments L0 to L1. The information collected within segments L0 to L1 is then sent to the analysis software. The analysis software analyzes the information within segments L1 to L2 to generate the first rule (i.e., the L0-L1 rule). Hardware operation within segments L1 to L2 is still based on the default rules. Within segments L2 to L3, the hardware operates based on the L0-L1 rules (because these rules are optimized, hardware operation in this segment is accelerated, and is referred to as accelerated operation based on the L0-L1 rules). Key information generated in segments L2 to L3 is collected and sent to analysis software. The analysis software can analyze the key information collected in segments L2 to L3 to generate new rules (which can be called the second rule), or it can analyze the key information collected in segments L0 to L3 to generate new rules (which can be called the third rule). The hardware still operates based on the first rule in segments L2 to L3, and based on the second or third rule in segments L3 to L4, and so on, in a loop, until the generated rules meet expectations.

[0080] Figure 6 is a second interactive timing diagram according to an embodiment of this application. Figure 6 shows the timing diagram of "pipeline processing". As shown in Figure 6, the on-chip system runs the target software in L0 to L1, collects the information generated in L0 to L1, and the analysis software analyzes the hardware information in L1 to L2 to obtain rules. At the same time, the on-chip system still runs based on the default rules in L1 to L2. In L2 to L3, the on-chip system accelerates its operation based on the rules generated in L0 to L1. In L2 to L3, the analysis software processes the information in L0 to L2 to obtain the running rules. In L3 to L4, the on-chip system accelerates its operation based on the rules generated in L0 to L2. In L3 to L4, the analysis software processes the information in L0 to L3 to obtain new running rules. After L4, the on-chip system accelerates its operation based on the rules in L0 to L3. After L4, the analysis software processes the information generated in L0 to L4 to obtain rules, and so on, until the obtained running rules meet the expectations.

[0081] Figure 7 is a schematic diagram of the interaction timing according to an embodiment of this application. Figure 7 shows the timing diagram of "one-time processing". As shown in Figure 7, the on-chip system runs the target software in the L0 to L1 segment, collects information in L1, and then the analysis software processes the key information in the L0 to L1 segment to obtain the running rules. After L2, the on-chip system accelerates the operation based on the L0-L1 rules. In Figure 7, only one rule is generated, hence the term "one-time processing".

[0082] In summary, in Figure 5, the hardware runs continuously. When it collects key information or receives data requests from the software, it generates key information data and transmits it to the software for processing. After calculating and generating new operating rules, the rules are sent to the hardware to guide its subsequent execution and prediction behaviors. Hardware operation and software processing are asynchronous; software processing can be executed intermittently, sending results to the hardware as soon as they are generated. During this period, the hardware continues to run acceleration programs. Essentially, new rules will improve hardware efficiency, but if no new rules arrive, it will still execute according to existing rules without blocking hardware operation. In Figure 6, hardware and software still execute asynchronously, but the system now performs highly efficient pipelined execution. While the hardware continues to run, the software also executes in real time at full bandwidth. The software continuously requests or reads hardware information and updates rules to the hardware frequently. This design is suitable for scenarios where dedicated hardware-accelerated programs change frequently. In Figure 7, offline processing software only needs to process hardware information once over a longer period and configure it for the hardware. This design is suitable for scenarios where the same type of dedicated program runs for a long time.

[0083] Regardless of the timing processing, Figures 5 through 7 all demonstrate similar interaction steps. Figure 8 is an interaction flowchart according to an embodiment of this application. As shown in Figure 8, the interaction flow includes the following steps: First, the program to be accelerated (i.e., the target software) is started. After the target software starts, it runs on the chip. During the running process, information (i.e., key information) generated by the chip when running the target software is collected. This collection process can also begin after receiving a request message from the analysis software. It should be noted that the chip has an internal cache used to store this key information. However, the cache size is limited. Therefore, it is necessary to retrieve the key information in a timely manner and save it in memory or other storage space. This allows for the storage of a relatively large amount of key information. Then, the collected key information is analyzed by the analysis software to generate rules. Since these rules guide the operation of the chip's hardware, they are also called hardware guidance rules. These hardware guidance rules are sent to the chip and used by the on-chip system when running the target software. In simple terms, Figure 8 illustrates the following steps: 1. Start; 2. The program to be accelerated starts; 3. The chip hardware runs the program to be accelerated; 4. The hardware generates key information, and if generated, it is transmitted to the software; 5. The information processing software program (i.e., analysis software) is run; 6. The analysis software generates rules and writes them to the chip's on-chip system to achieve acceleration.

[0084] Figure 9 is a schematic diagram of the interaction flow according to an embodiment of this application. As shown in Figure 9, the analysis software analyzes information from the system-on-a-chip (SoC) through software analysis algorithms. Since the information obtained from the SoC is acquired based on different target software, the rules generated by different target software may be different. Therefore, corresponding rules can be generated for the information generated by the operation of different target software on the SoC, which is a customized processing for the target software. The SoC includes a central processing unit (CPU) and a cache. The CPU is used to execute the target software, and the cache is used to store the instructions and data corresponding to the target software. Then, the SoC generates information when executing the target software (the so-called real-time information needs to be continuously retrieved from the cache and stored in a large storage space such as memory). After the analysis software obtains rules based on the information, the CPU can determine the operations that need to be performed in the future based on the rules, thereby performing advance processing (therefore, the rules can also be understood as a prediction of future behavior), which can improve the efficiency of the SoC. In Figure 9, the central processing unit executes the program to be accelerated. The system-on-chip (SoC) stores the key information generated in real time (i.e., real-time information) in memory. The analysis software develops customized analysis algorithms based on different dedicated programs. The input of the algorithm is the real-time information generated by the hardware. The software generates future trajectory guidance rules required for hardware execution and simultaneously issues rules that can locate execution rules. When the SoC executes the target software, it searches for rules in memory addresses. The search rules have higher information density and occupy less area, which can reduce search time complexity and avoid bandwidth loss and increased search latency caused by directly reading execution rules. The SoC issues search requests and generates the information required for the search based on the accuracy of internal predictions. The search information and search rules are matched and searched to find the fragment address pointer of the corresponding rule. After reading the hardware execution rule, its information is extracted, and the prediction information is input to the SoC for subsequent prediction behavior.

[0085] In the above implementation, an interface can be pre-configured in the processor (or chip) to obtain information generated during the execution of the target software. That is, the hardware exposes an information interface, and the software reads the information and processes it offline. Leveraging the flexibility of the software, specific optimizations can be implemented for different programs to achieve optimal performance for each program. Traditional processors do not perform specific optimizations for each program, or rather, the processor itself cannot be specifically optimized. Software engineers can only adapt the program or compiler to the processor to achieve better performance. In the above implementation, the processor can be adapted to the program without changing the hardware, as modifying the hardware is very costly. However, by fully utilizing the flexibility of the software processing algorithm, different rules can be designed according to the running program (these rules can be understood as rules specific to the target software or program), generating program-specific rules to guide hardware operation and improve performance.

[0086] In one alternative implementation, the software-side offline algorithm analyzes and processes real-time information provided by the processor hardware. The processed rules can be time-independent discrete fragments or long historical fragments with strong temporal and spatial locality. These rules can predict future processor behavior, but to avoid consuming too much storage space, compression algorithms can be used to compress the rules, thus storing more rules in less storage space, achieving a predictive architecture design with long historical information and guidance information.

[0087] In the above implementation, a software-hardware co-design, such as offline software analysis, is employed. This requires information generated by the target program running on the on-chip system (this information is used for future predictions and is therefore also called historical information). Since the processor typically only stores historical information in on-chip cache (RAM or SRAM), the amount of historical information is limited, making it impossible to perform long-term historical analysis and processing. For example, dynamic branch prediction algorithms such as Branch History Table (BHT), GShare, and Tagged Geometric History (TAGE) primarily store historical information in SRAM. The size of SRAM determines the length of historical information that can be stored, so the traditional approach is to increase the processor area to improve performance. The above implementation utilizes on-chip or off-chip memory to store long-term historical information as a prediction data source. The processor hardware exposes an information interface to the offline analysis software. The offline analysis software leverages its flexibility to perform specific analysis on the dedicated program (i.e., the target software), providing rule guidance for the future operation of the on-chip system. The hardware then performs high-performance prediction and execution based on these rules.

[0088] Through the above implementation methods, offline software utilizes long-term historical information provided by the hardware for analysis and processing, ensuring prediction accuracy and thereby improving processor execution performance. These embodiments guarantee the flexibility of the hardware and software system without requiring hardware modifications, thus saving costs.

[0089] The above implementation methods, based on hardware information upload and offline software processing to provide hardware execution rules, can accelerate the following aspects: branch prediction; i-cache and d-cache prefetching, improving cache hit rate; TLB prefetching, improving TLB hit rate, and avoiding direct page table reading. In these implementation methods, by predicting the time nodes of multi-core data updates and the data itself, the performance and power consumption losses caused by repeated updates due to the broadcast mechanism in traditional multi-core cache consistency are resolved. Obtaining the future prediction trajectory can guide the prefetch behavior in the local host cache of the distributed cluster, thereby effectively reducing the impact of network latency on system performance and improving performance.

[0090] The computer program corresponding to the aforementioned analysis software can also be loaded onto a computer or other programmable data processing device, causing a series of operational steps to be executed on the computer or other programmable device to produce computer-implemented processing. The instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams. Different steps can be implemented by different modules. This embodiment provides such a device or system. This system or device is used to implement the functions of the methods in the above embodiments. Each module in the system or device corresponds to each step in the method, as already described in the method, and will not be repeated here.

[0091] In one embodiment, a chip is also provided, including a central processing unit and a cache; wherein the system-on-a-chip on the chip is configured to perform the following method steps: acquiring rules, wherein the rules reflect a prediction of the behavior of the system-on-a-chip when running target software, the rules being obtained by analyzing runtime information collected from the historical runtime of the system-on-a-chip when running the target software, the runtime information being generated by the system-on-a-chip running the target software; and taking measures in advance according to the rules when running the target software, wherein the measures are determined according to the rules, the measures taken in advance predicting future operations of the system-on-a-chip and preparing for the operations.

[0092] As an optional implementation, it further includes: an interface, wherein there are one or more interfaces, the interfaces being used to send the operating information to an external part of the system-on-a-chip, and to receive rules from the analysis software or to receive parameters from the analysis software; wherein the parameters are used in an algorithm running on the system-on-a-chip to generate the rules, and the algorithm is used to generate the rules.

[0093] In one embodiment, a chip optimization processing system is also provided for the aforementioned chip. This system includes the chip and analysis software. The analysis software performs the following steps: receiving runtime information, wherein the runtime information is obtained through an interface pre-configured on the chip, and after being read from the interface, is stored in storage space outside the chip; analyzing the runtime information to determine the rule; sending the rule to the on-chip system; or, obtaining parameters for generating the rule based on the runtime information and sending the parameters to the on-chip system; wherein an algorithm running on the on-chip system uses the parameters to generate the rule; the algorithm is used to generate the rule. Other steps performed by the analysis software have been described in the above embodiments and will not be repeated here. For example, the computing resources running the analysis software are different from the computing resources of the on-chip system running the target software.

[0094] Optionally, when the chip includes multiple cores, each core comprises two parts: a first part for running the analysis software and a second part for running the target software as a system-on-a-chip; or, the analysis software runs on a computing device different from the computing device on which the chip is located; or, the analysis software runs on the system-on-a-chip, which uses different time-domain resources to run the analysis software and the target software in a time-sharing manner.

[0095] The above implementation method solves the problems of not being able to cover other software and high cost caused by using hardware accelerators when optimizing processors for specific software, thereby enabling processor performance optimization for different target programs.

[0096] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A chip optimization processing method, characterized in that, include: On-chip system acquisition rules located on the chip, wherein the rules reflect the behavior prediction of the on-chip system when running target software, and the rules are obtained by analyzing the running information collected when the on-chip system runs the target software in the past, the running information being generated by the on-chip system running the target software; When the on-chip system runs the target software, it takes measures in advance according to the rules. The measures are determined according to the rules, and the measures taken in advance estimate the operations that the on-chip system will perform in the future and prepare for the operations.

2. The method according to claim 1, characterized in that, It also includes: analysis software receiving the operational information, wherein the operational information is obtained through an interface pre-configured on the chip, and after being read from the interface, the operational information is stored in a storage space located outside the chip; the analysis software analyzes the operational information and determines the rule; the analysis software sends the rule to the system-on-a-chip; or, The on-chip system acquires the rules by: the analysis software receiving the runtime information; the analysis software obtaining parameters for generating the rules based on the runtime information; the on-chip system receiving the parameters and using the parameters in an algorithm running on the on-chip system to generate the rules; wherein, the algorithm is used to generate the rules.

3. The method according to claim 2, characterized in that, The computing resources used to run the analysis software are different from those used to run the target software on the on-chip system.

4. The method according to claim 3, characterized in that, In the case where the chip includes multiple cores, each core comprises two parts: a first part runs the analysis software, and a second part runs the target software as a system-on-a-chip; or... The analysis software runs on a computing device, which is different from the computing device where the chip is located; or... The analysis software runs on the on-chip system, and the on-chip system uses different time-domain resources to run the analysis software and the target software in a time-sharing manner.

5. The method according to claim 4, characterized in that, The target software runs on multiple chips, which are located on different computing devices. The analysis software receives runtime information generated by the on-chip systems running the target software on each of the multiple chips. The analysis software analyzes the runtime information generated by all the on-chip systems to obtain the rules. The analysis software transmits the rules to the on-chip systems on each of the multiple chips.

6. A chip comprising a central processing unit and a cache; wherein, The on-chip system is used to perform the following method steps: The rules are obtained, wherein the rules reflect the prediction of the behavior of the on-chip system when running the target software, and the rules are obtained by analyzing the running information collected when the on-chip system runs the target software in the past, and the running information is generated by the on-chip system running the target software; When the target software is running, measures are taken in advance according to the rules, wherein the measures are determined according to the rules, and the measures taken in advance predict the operations that the on-chip system will perform in the future and prepare for the operations.

7. The chip according to claim 6, characterized in that, It also includes: an interface, wherein there are one or more interfaces, the interfaces being used to send the operating information to an external part of the on-chip system, and to receive rules from the analysis software or to receive parameters from the analysis software; wherein the parameters are used in an algorithm running on the on-chip system to generate the rules, and the algorithm is used to generate the rules.

8. A chip optimization processing system, characterized in that, Including the chip and analysis software as described in claim 6 or 7, wherein the analysis software is used to perform the following method steps: The system receives operational information, which is obtained through an interface pre-configured on the chip. After being read from the interface, the operational information is stored in a storage space located outside the chip. The system analyzes the operational information to determine the rule and sends the rule to the system-on-a-chip. Alternatively, the system obtains parameters for generating the rule based on the operational information and sends the parameters to the system-on-a-chip. The algorithm running on the system-on-a-chip uses the parameters to generate the rule.

9. The system according to claim 8, characterized in that, The computing resources used to run the analysis software are different from those used to run the target software on the on-chip system.

10. The system according to claim 9, characterized in that, In the case where the chip includes multiple cores, each core comprises two parts: a first part runs the analysis software, and a second part runs the target software as a system-on-a-chip; or... The analysis software runs on a computing device, which is different from the computing device where the chip is located; or... The analysis software runs on the on-chip system, and the on-chip system uses different time-domain resources to run the analysis software and the target software in a time-sharing manner.